Tech

Anthropic details technical implementation of Claude AI watermarks to meet EU regulations

As the EU AI Act’s Transparency Code takes effect, Anthropic outlines how undetectable patterns will be embedded in low-stakes word choices, distinguishing its method from stylistic detection services.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · View original source
Anthropic shares more details about how Claude’s new watermarks will work
Company adopts Google DeepMind’s SynthID-Text approach, addressing user concerns over privacy and output integrity

Anthropic has published a detailed technical breakdown of the watermarking system being implemented in its Claude chatbot, a move designed to comply with the European Union’s AI Act Transparency Code. The company confirmed it will utilise the SynthID-Text approach originally outlined by Google DeepMind in 2024, embedding undetectable patterns into the model’s responses that can be identified via a forthcoming detection API.

The implementation relies on creating subtle patterns during "low-stakes choices," such as selecting between synonyms like "overcast" and "grey" to describe weather conditions. Anthropic asserts that these patterns remain invisible to human readers and do not degrade the quality or utility of the generated text. The company distinguishes this method from AI detection services like Pangram, which analyse writing for stylistic "tells" or structural patterns, noting that checking for embedded watermarks is fundamentally different from identifying linguistic habits.

This technical disclosure follows significant user backlash and reported subscription cancellations on social media platforms. While anecdotal reports from Business Insider suggest dozens of users have cancelled their subscriptions on X, and discussions on Reddit have ranged from conspiracy theories to accusations of deceptive practices, Anthropic maintains that the watermarking is a necessary compliance measure. The company noted that other major model developers have signed the same Code of Practice and will similarly implement their own watermarks.

Regarding the resilience of the watermarks against editing, Anthropic stated that light editing is unlikely to remove the pattern completely. However, a complete rewrite where every word is replaced would effectively eliminate the watermark. The company also clarified that for text lightly proofread or edited by Claude, the watermark’s presence depends on the length of the text and the extent of the AI’s involvement; if a human author writes nearly all the words, there is little for the watermark to attach to.

In the specific context of code generation, Anthropic explained that watermarks will have a negligible effect on the actual code produced. Because the model must prioritise functional syntax over arbitrary word choices, there is limited freedom to embed patterns. However, watermarks can still be applied to comments within code where arbitrary terminology choices exist. The company plans to release a detection API to allow third parties to verify the presence of these watermarks in the future.

Continue reading

More from Tech

Read next: Tiny386 PC emulator ported to Raspberry Pi Pico 2-class hardware
Read next: Readers turn to AI chatbots for personalised fiction and role-play
Read next: Septuagint’s contested history comes into focus in review of Timothy Michael Law’s book