AI watermarking could make LLM guardrail adherence unpredictable — and that could be a big problem for the EU AI Act
AI watermarking aims to prove the authenticity of any text New study finds it also changes LLM behavior – and in a bad way EU AI Act could mean more models have watermarks despite side effects New Lasso research has revealed that AI watermarking could actually unintentionally change how LLMs behave
<![CDATA[ <article> <ul><li><strong>AI watermarking aims to prove the authenticity of any text</strong></li><li><strong>New study finds it also changes LLM behavior – and in a bad way</strong></li><li><strong>EU AI Act could mean more models have watermarks despite side effects</strong></li></ul><p>New Lasso <a href="https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-on-ai-agent-behavior" target="_blank" rel="nofollow">research</a> has revealed that AI watermarking could actually unintentionally change how LLMs behave following the testing of Google DeepMind's SynthID-Text.</p><p>The company's researchers found that SynthID-Text can change whether models refuse harmful requests, their susceptibility to prompt injection, which tools AI agent choose and more.</p><p>However, at its core, SynthID-Text and other similar watermarking is only designed to hide a machine-readable indicator as to whether text was AI-generated or human-written.</p><h2 id="researchers-find-that-ai-watermarking-can-unintentionally-change-ai-behavior">Researchers find that AI watermarking can unintentionally change AI behavior</h2><p>The "watermarking procedure can therefore affect both what the model says and what an agent does," Lasso concludes, referring to the side effect as "sampling drift."</p><p>One of the biggest concerns highlighted by the paper is that, even without an attack, watermarking changed some of the models' refusal decisions, making them more willing to answer potentially harmful prompts. Combined with prompt injection, Lasso found the consequences more amplified.</p><p>Despite the unintended consequences, Anthropic recently <a href="https://www.anthropic.com/news/claude-text-watermark" target="_blank" rel="nofollow">announced</a> that future generations of Claude would use AI watermarking similar to Google DeepMind's, stressing that one of the key drivers was to adhere to the EU AI Act. With that in mind, AI watermarking is set to become far more mainstream across other model providers, making these mishaps far more common and leading to further security concerns.</p><p>Ultimately, Lasso urges developers to rerun benchmarks, safety evaluations and other tests to check for any unintended consequences, rather than just applying it blindly to existing configurations.</p><p>"These findings make reassessment important whenever watermarking is introduced or its configuration or key changes," Lasso concludes, stressing that the work shouldn't be taken as an argument against AI watermarking for provenance.</p><p>Additionally, the research presents a new angle on AI watermarking, because until now researchers have largely focused on whether watermarks can be both applied and detected effectively. Few have uncovered such security-focused consequences as this one.</p><figure class="van-image-figure pull-right inline-layout" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' style="max-width:676px;"><p class="vanilla-image-block" style="padding-top:31.51%;"><img id="diM9tpwF2Lz85R8q85CT78" name="tr-g_news" alt="Google logo on a black background next to text reading 'Click to follow TechRadar'" src="https://cdn.mos.cms.futurecdn.net/diM9tpwF2Lz85R8q85CT78-1920-80.jpg" mos="" align="right" fullscreen="" width="676" height="213" attribution="" endorsement="" class="pull-rightinline"></p></div></div></figure> </article> ]]>
Read the full article on TechRadar
Read Full Article →