AI watermark tests find shifts in agent calls and refusals
Lasso Research found SynthID-Text changed tool-call accuracy and, on some models, weakened refusal under prompt injection. Results varied by model and watermark key.
Watermarks designed to identify AI-generated text may also affect how language models behave, according to benchmark tests by Lasso Research. The study tested SynthID-Text across six models and found changes in both tool calling and refusal behaviour.
Watermarking reduced tool-call accuracy on six of seven tested models, the researchers reported. Paired runs also showed that individual decisions changed more often than aggregate accuracy alone suggested: across 21 model and temperature combinations, an average 6.5% of call verdicts differed between watermarked and unwatermarked runs.
In prompt-injection tests, watermarking weakened refusal on several models, though effects depended on the model and key. For Gemma 3 27B, the share of harmful prompts where the verdict changed rose from 6.0% without injection to 23.5% with a fixed injection technique. The net compliance change shifted from minus 1.0 to plus 12.5 percentage points.
The study measured refusal at model level, not end-to-end agent actions, and did not test the combined case of weakened refusal and tool use. Its findings come from benchmark experiments, not reported incidents in deployed systems. The injection test also used one fixed technique.
SynthID-Text is a generation-time watermark: it can change token selection under a fixed key, even though its non-distortionary guarantee applies over watermark randomness. The researchers said effects varied with the key as well as the model.
Lasso Research recommends evaluating and red-teaming agents with the watermark configuration intended for deployment, including under prompt injection. The study discusses watermarking amid regulatory interest in machine-readable marking of AI-generated text, but does not establish whether any deployment meets regulatory requirements.


