A security research report published this fall is forcing AI labs to confront an uncomfortable tradeoff: the same invisible watermarks regulators are demanding to flag AI-generated text can also change how AI agents behave. According to research from Lasso Security, AI watermarking reduced tool-calling accuracy in six of seven tested models and made several models more likely to comply with harmful requests when the watermark was combined with a prompt injection attack. The report, written by Lasso researcher Andrea Siposova and first covered by The Register’s Thomas Claburn, lands right as the EU AI Act’s content-marking rules move from theory to enforcement.
The timing matters. Anthropic has said future Claude models will embed an invisible watermark in their output, built on Google DeepMind’s SynthID-Text method, which The Register reports has also been adopted by OpenAI. Both companies are responding to Article 50(2) of the EU AI Act, which requires providers of AI systems that generate synthetic text, audio, image, or video to mark that output in a machine-readable, detectable format. Lasso’s findings suggest that compliance has a side effect nobody asked for: it can quietly change what an AI agent decides to do, especially under attack.
What Lasso Security Actually Found
Lasso’s study, titled “The Provenance Tax,” ran a paired experiment across seven open-weight models: Phi-4, Llama-3.1-8B, Qwen3-32B, Qwen3-4B, Gemma-3-12B, Gemma-3-27B, and Granite-3.2-8B. Each model generated the same prompts twice, once with Google’s SynthID-Text watermarking turned on and once with it off, using identical seeds, batch order, and temperature settings. The only variable that changed was the watermark itself.
The researchers call the resulting behavior change “sampling drift.” SynthID-Text works by nudging the model’s token-selection process using what’s known as tournament sampling, biasing low-stakes word choices just enough to leave a statistical fingerprint. In ordinary prose, that bias is invisible to a reader. But inside a structured output, like a tool call’s arguments, those same low-stakes tokens can be the difference between a correct function call and one that silently does the wrong thing. As Lasso put it, “at the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection,” a point made directly in the company’s research write-up.
How the Test Was Built: BFCL, HarmBench, and a Fixed Injection
Lasso split the work into two tracks. Tool-calling accuracy was measured against BFCL v4 single-turn AST, a widely used function-calling benchmark, covering both live and non-live call-expected tasks. Refusal behavior was tested on 200 harmful behaviors from HarmBench plus 100 benign controls from JailbreakBench, run both as bare requests and under one fixed prompt injection technique that claims the safety filter has been disabled and demands compliance.
That prompt injection design mirrors a threat the industry has already flagged elsewhere this year. The BIPIA and Deepset benchmarks used to stress-test models like Security-One 27B measure similar indirect injection attacks, and researchers testing Copilot CLI, Grok, and Gemini found comparable agent hijacking paths earlier this year. What’s new in Lasso’s work is isolating watermarking itself as a variable that shifts how often those attacks succeed, independent of any flaw in the underlying model’s training.
Lasso ran its watermarking layer using HuggingFace’s unmodified SynthIDTextWatermarkLogitsProcessor, with the following configuration:
SynthIDTextWatermarkLogitsProcessor(
tournament_layers=30,
ngram_len=5,
sampling_table_size=2**16,
context_history_size=1024
)
That is the same non-distortionary configuration described in Google DeepMind’s prior SynthID research (credited in Lasso’s report to Dathathri and colleagues), meaning the watermark is designed, in theory, to leave the model’s overall output distribution unchanged. Lasso’s point is that “non-distortionary” on paper does not mean behaviorally invisible in practice, particularly once a fixed watermark key enters the picture.
Tool-Calling Accuracy Drops in 6 of 7 Models
On the core tool-calling test, watermarking reduced accuracy on six of the seven models, with what Lasso describes as a significant decrease on four of them. But the aggregate accuracy number undersells what’s happening. Lasso also measured “churn,” the paired disagreement rate between a model’s watermarked and unwatermarked answers on the exact same question, and found churn consistently higher than the net accuracy change alone would suggest.
At a temperature setting of 1.0, Phi-4’s net accuracy loss was 2.87 points, but 16.8% of its individual tool-call verdicts flipped between the watermarked and unwatermarked runs, meaning some calls got worse while others got better and mostly canceled out in the aggregate score. Llama-3.1-8B showed the same pattern at a smaller scale: a 0.87-point net loss masked a 9.9% churn rate. Across all 21 model-temperature combinations Lasso tested, churn averaged 6.5%, and the statistical confidence interval excluded zero in every single case, meaning the effect was not noise.
| Metric | Result |
|---|---|
| Models tested (tool calling, BFCL v4) | 7: Phi-4, Llama-3.1-8B, Qwen3-32B, Qwen3-4B, Gemma-3-12B, Gemma-3-27B, Granite-3.2-8B |
| Models with reduced accuracy under watermarking | 6 of 7 |
| Models with a “significant” accuracy decrease | 4 of 7 |
| Phi-4 net accuracy change (T=1.0) | -2.87 points |
| Phi-4 paired churn rate (T=1.0) | 16.8% |
| Llama-3.1-8B net accuracy change (T=1.0) | -0.87 points |
| Llama-3.1-8B paired churn rate (T=1.0) | 9.9% |
| Average churn across 21 model-temperature combinations | 6.5% |
| Dominant error type on Phi-4 / Granite-3.2-8B | Malformed output (-5.96 / -4.36 points) |
| Dominant error type on Llama-3.1-8B | Wrong arguments (-3.48 points) |
Source: Lasso Security, “The Provenance Tax,” September 2026.
Where the Errors Come From: Wrong Tools, Bad Arguments, Broken Syntax
The failure modes vary by model, and that matters for anyone trying to defend against this. On Llama-3.1-8B, most of the accuracy loss came from incorrect arguments, a wrong path, recipient, or amount passed to an otherwise correctly chosen tool, which is arguably the most dangerous failure because the call still executes. On Phi-4 and Granite-3.2-8B, malformed output dominated instead: watermarking pushed enough token-level noise into structured fields like JSON braces and function names that the call failed to parse at all.
That distinction echoes a broader pattern security teams have been tracking all year in agentic AI deployments. SalesBleed’s three flaws in Salesforce’s Agentforce showed how a technically well-formed but subtly wrong tool call can be more damaging than an obvious crash, because nothing alerts a human that anything went wrong. Lasso’s research adds a new wrinkle: the same failure can now be introduced not by a prompt injection attacker, but by a compliance feature the provider added voluntarily.
Prompt Injection Makes the Safety Problem Worse
The refusal-behavior results are where Lasso’s findings turn from a quality-control issue into a security one. On bare harmful requests, with no attack involved, watermarking’s effect on refusal was mostly small. But layer in the fixed prompt injection technique, and the picture changes sharply.
For Gemma-3-27b, churn under injection jumped from 6.0% to 23.5%, and the net compliance shift flipped from -1.0 points to +12.5 points, meaning the model became measurably more willing to comply with a harmful request once both watermarking and prompt injection were present together. Gemma-3-12b showed the same direction of effect at a smaller scale: churn rose from 7.5% to 11.0%, with net compliance moving from -0.5 to +9.0 points. Llama-3.1-8B’s churn under injection reached 14.0% at a low temperature and 17.5% at a higher one. As Lasso summarized it, “watermarking changes refusal behavior on bare harmful requests, but the effect becomes more pronounced under prompt injection,” a finding laid out in the company’s published report.
Not every model moved. Phi-4 and Qwen3-4B barely budged under either condition, but Lasso cautions that’s likely because both models already over-refuse by default, including on harmless control prompts, so there wasn’t much room left for the effect to show up.
Watermark Churn vs. Temperature Churn
To put the watermark effect in perspective, Lasso compared it against a known source of model variability: raising the sampling temperature. The watermark-induced churn under prompt injection turned out to be significantly higher than temperature-induced churn on four of the six models tested in this part of the study.
| Model | Watermark Churn (T=0.7) | Temperature Churn (0.001→0.7) | Difference |
|---|---|---|---|
| Gemma-3-27b | 26.0% | 13.5% | +12.5 pts |
| Granite-3.2-8B | 21.5% | 15.5% | +6.0 pts |
| Llama-3.1-8B | 17.5% | 7.5% | +10.0 pts |
| Gemma-3-12b | 11.0% | 6.0% | +5.0 pts |
| Phi-4 | 0.5% | 0.5% | +0.0 pts |
| Qwen3-4B | 0.0% | 0.0% | +0.0 pts |
Source: Lasso Security, “The Provenance Tax,” refusal churn under prompt injection at T=0.7, compared against temperature-induced churn with watermarking off.
In other words, on four of six models, simply turning on AI watermarking changed safety-relevant behavior under attack by more than doubling the sampling temperature would have. That’s a bigger behavioral swing than most red teams are currently accounting for when they certify a model for agentic deployment.
The Watermark Key Problem
Lasso also tested whether the effect depends on which specific watermark key is used, since SynthID-Text’s output varies with the key even though the underlying method stays the same. Running eleven different keys at a 0.7 temperature setting, the team found the effect swings substantially depending on which key is active. For Llama-3.1-8B, the key used in the main study increased attack success by 3.5 points, while the other ten keys averaged a 4.4-point increase, ranging as high as 14.5 points and as low as a 4.5-point decrease.
That variability is a problem for anyone trying to certify a model once and assume it stays certified. If a cloud provider or model vendor rotates its watermark key, which Lasso notes may happen “outside the agent developer’s direct control” when the provider hosts the model, the safety profile of every downstream agent built on that model can shift without any code change on the developer’s end.
Why Anthropic and OpenAI Are Watermarking Now
None of this is happening in a vacuum. Anthropic has announced that future Claude models will embed an invisible watermark built on SynthID-Text, and The Register reports OpenAI has adopted the same underlying Google DeepMind method. Neither company is doing this purely out of goodwill: Article 50(2) of the EU AI Act now requires providers of AI systems that generate synthetic audio, image, video, or text to mark that content in a machine-readable, detectable format.
The European Commission has been explicit about the scope of that obligation. Per its official guidance, “providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, must ensure that AI-generated or manipulated content are marked in a machine-readable format and detectable as artificially generated or manipulated,” as stated in the Commission’s Article 50 FAQ. A separate Commission fact page adds that “providers must apply a machine-readable mark to synthetic content generated or manipulated by AI and enable its detection, unless the AI system performs an assistive function for standard editing or does not substantially alter the input data,” according to the Commission’s quick-facts page on transparency rules.
The EU AI Act Deadline Behind the Rush
Article 50’s marking obligation applies from August 2, 2026, for AI systems placed on the market from that date onward. Systems that were already on the market before August 2, 2026, get a transition window, with the obligation kicking in on December 2, 2026. That means every major provider serving EU customers is either already marking synthetic text output or racing to do so before the December deadline lands.
That compressed timeline is part of why Lasso’s findings landed as uncomfortably as they did. Watermarking was framed, correctly, as a transparency and provenance tool, a way to help platforms and regulators tell real content from AI-generated content. Nobody designing that policy was thinking about what happens when the same text becomes the reasoning trace an autonomous agent uses to decide which tool to call next. Lasso’s research is the first widely reported attempt to measure that gap directly, rather than just flag it as a theoretical concern.
Market Impact: What This Means for Enterprises Running AI Agents
For enterprises deploying AI agents against regulated European users, this creates a genuine bind. Skipping watermarking isn’t an option once Article 50 fully applies, but Lasso’s data shows that turning it on can measurably change tool-calling accuracy and, under attack conditions, safety-refusal behavior. That lands at an awkward moment: regulators are already circling agentic AI deployments more broadly, with the FTC’s inquiry into OpenAI and Anthropic over AI agent attacks underscoring how closely agent behavior is now being scrutinized on both sides of the Atlantic.
Security teams evaluating agent deployments now have one more variable to control for. Model cards and provider benchmark disclosures typically describe performance with no mention of whether watermarking was active during testing, and Lasso’s results suggest that omission matters more than most vendors have treated it. The pattern lines up with other recent findings about how fragile AI agent safety guarantees can be in practice, from GLM-5.3’s reported 100% safety bypass rate under adversarial testing to the broader push behind Nvidia’s agent safety platform, which has signed up more than 100 partners specifically to address gaps like this one.
For vendors, the near-term fix Lasso recommends isn’t abandoning watermarking, it’s treating it as part of the deployed security configuration rather than a bolt-on compliance checkbox. That means re-running red-team evaluations, including prompt injection tests, under the exact watermark settings a production agent will actually use, not just under the clean conditions most benchmark leaderboards report.
Historical Context: From Content Provenance to Agent Safety
Text watermarking itself isn’t new. Google DeepMind researchers described the “non-distortionary” approach Lasso tested, crediting prior work by Dathathri and colleagues, years before agentic AI deployment became mainstream. Image and audio watermarking for AI-generated media has existed even longer, largely aimed at combating deepfakes and misinformation. What’s changed is the application layer sitting on top of the model: when an LLM’s output was just text a human read, a watermark’s lexical nudges were a non-issue. Now that the same output routes directly into function calls, file operations, and autonomous decisions, those nudges become operational risk.
This is the same structural shift that has driven a wave of prompt injection research across the industry this year, from benchmark efforts like BIPIA and Deepset to live exploits against production coding assistants. Lasso’s contribution is narrower but concrete: it shows that even a compliance-driven, provider-side feature, not an attacker-controlled input, can shift an agent’s safety behavior in a measurable, reproducible way.
Competitive Comparison: How Major AI Providers Are Positioned
Public disclosure about watermarking deployment remains limited, but the available reporting draws a rough map. Google DeepMind originated SynthID-Text and continues to publish the underlying research. Anthropic has publicly committed to embedding the same watermarking approach in future Claude models, applied at the model level across the Claude Platform API and supported cloud providers, according to Lasso’s report. The Register additionally reports that OpenAI has adopted SynthID-Text as well. Other major model providers have not been confirmed in current reporting as having deployed comparable watermarking at this scale, which means Lasso’s findings, drawn from smaller open-weight models like Phi-4, Llama, Qwen3, Gemma, and Granite, may not map directly onto the frontier models EU enterprises are most likely to deploy in agentic settings.
| Provider | Watermarking Status (per current reporting) |
|---|---|
| Google DeepMind | Originated SynthID-Text; publishes underlying research |
| Anthropic | Announced invisible watermark for future Claude models, based on SynthID-Text |
| OpenAI | Adopted SynthID-Text, per The Register’s reporting |
| Open-weight models tested by Lasso (Phi-4, Llama-3.1-8B, Qwen3, Gemma-3, Granite-3.2-8B) | Watermarking applied experimentally in Lasso’s study, not a vendor production deployment |
That gap is worth watching. If frontier, closed-weight models respond to SynthID-Text differently than the open-weight models Lasso tested, either better or worse, enterprises won’t know which way that cuts until someone runs the same paired-evaluation methodology against GPT and Claude model families directly.
What Security Teams Should Do Before the December Deadline
Lasso’s recommendation, echoed by the OWASP GenAI LLM Top 10 2026’s treatment of prompt injection as an input-side vulnerability with consequences that extend to unauthorized tool actions, is straightforward in principle and harder in practice: don’t evaluate agent safety once and assume it holds. Any change to watermark configuration, including a routine key rotation initiated by a model provider, can shift tool-calling accuracy and refusal behavior in ways that aggregate benchmark scores will not reveal. Paired testing, the same prompt run with and without the production watermark setting, is the only way Lasso’s data shows this effect reliably.
That’s a meaningfully higher evaluation burden than most teams currently budget for, and it arrives at the same time agentic AI deployments are already under regulatory pressure, including the scrutiny AI agents have drawn following incidents covered in the 28-second hijack of Copilot CLI, Grok, and Gemini earlier this year.
What Happens Next: Five Predictions
- More independent security vendors will publish their own paired watermark-on/watermark-off evaluations before the EU’s December 2, 2026 compliance deadline, turning this into a standard red-team category rather than a one-off study.
- Expect pressure on Anthropic, OpenAI, and Google to disclose whether their production watermarking configurations were tested against the exact failure modes Lasso identified, particularly wrong-argument tool calls and refusal flips under injection.
- Enterprise AI agent platforms will likely start offering a documented option to evaluate, and in some deployments disable, watermarking for internal, non-EU-facing agentic workloads where Article 50 doesn’t apply.
- Watermark key rotation policies will become a disclosure point in vendor security questionnaires, since Lasso’s key-sensitivity results show the same watermark method can swing attack success by double digits depending on which key is active.
- Academic and industry researchers will push for watermarking designs that avoid touching decision-relevant tokens in structured output, separating the provenance signal from the tokens an agent actually acts on.
Frequently Asked Questions
What is AI watermarking and why are companies adding it now?
AI watermarking embeds a statistical, machine-readable pattern into AI-generated content, like text, images, or audio, so it can later be identified as AI-generated. Companies are rolling it out now largely because Article 50(2) of the EU AI Act requires providers of AI systems generating synthetic content to mark that output in a detectable, machine-readable format, with the obligation applying from August 2, 2026.
What did Lasso Security’s report actually find?
Lasso found that Google’s SynthID-Text watermarking method reduced tool-calling accuracy in six of seven tested AI models on the BFCL v4 benchmark, and that under a combined prompt injection and harmful-request test, several models became measurably more likely to comply with requests they would otherwise refuse.
Which AI models were tested in the study?
Lasso tested seven open-weight models: Phi-4, Llama-3.1-8B, Qwen3-32B, Qwen3-4B, Gemma-3-12B, Gemma-3-27B, and Granite-3.2-8B. Six showed reduced tool-calling accuracy under watermarking, and six of those models (excluding Qwen3-32B) were included in the refusal and prompt-injection testing.
Does this mean AI watermarking should be turned off?
Lasso’s report explicitly stops short of that conclusion. The findings argue that watermarking needs to be evaluated as part of an agent’s deployed security configuration, with red-teaming repeated under the production watermark settings, rather than being treated as a compliance checkbox with no safety review.
Are Anthropic and OpenAI affected by these findings?
Lasso’s direct testing covered smaller open-weight models, not Anthropic’s or OpenAI’s production systems. However, The Register reports both companies have adopted SynthID-Text, the same underlying watermarking method Lasso tested, which means the sampling-drift mechanism it describes is architecturally relevant to their models even though per-model results have not been independently published.
When does the EU AI Act’s watermarking requirement take full effect?
Article 50(2) applies from August 2, 2026, for systems placed on the market from that date. Systems already on the market before that date get a transition period, with the obligation applying from December 2, 2026.
What is “churn” in Lasso’s methodology?
Churn, or paired disagreement rate, measures the share of individual test items where a model’s answer changed between the watermarked and unwatermarked runs. Lasso used this metric because net accuracy or refusal-rate changes can mask the real effect: a call that becomes wrong can be offset by another that becomes right, leaving the aggregate score nearly flat even though the model is behaving differently item by item.




