Demystifying Claude's New Watermarking System
Anthropic published a blog post Friday seeking to answer basic questions about how it will watermark text generated by its chatbot Claude. The company aimed to clarify how the system operates, whether users can hide it through editing, and how the changes affect programming code.
The system relies on "low-stakes choices" during text generation, such as selecting between similar descriptive words like "overcast" and "grey". By manipulating these minor vocabulary choices, Claude creates a specific linguistic pattern in its responses.
This pattern remains completely undetectable to a standard reader but is easily identified by anyone possessing the key that encodes it. Anthropic stressed that this process will not degrade performance, noting that watermarking does not impact the quality of Claude's output.
“To a reader, a watermarked response is indistinguishable from an unwatermarked one”. More specifically, Anthropic said it will be using the SynthID-Text approach that the Google DeepMind team outlined in 2024, and that it plans to release a watermark detection API.
Anthropic distinguished this cryptographic system from existing AI detection tools built by companies like Pangram, which search for writing "tells" like the phrasing "this isn't [X], it's [Y]". The company noted that picking up on conversational patterns is fundamentally different from checking for an embedded watermark.
Regarding attempts to bypass the tracking, Anthropic noted that light editing is unlikely to completely remove the watermark. While replacing every single word would eliminate the tracking, the company argued that after a complete rewrite, it is debatable whether the text can still be described as AI-generated.
For content that Claude merely proofreads, the watermark's presence depends on the length of the text and how heavily it was altered. In lightly edited documents, nearly all words remain human-written, leaving little for the watermark to attach to.
Generated programming code will also feature a significantly reduced watermark footprint. Because functional code restricts the model's freedom to choose between varied options, the watermark will primarily reside in arbitrary areas like code comments, resulting in a negligible effect on the actual code produced.
The Backlash and the EU AI Act
Claude users have been debating the move since the company revealed earlier this week that it would be doing this watermarking to comply with the EU AI Act’s Transparency Code. This legislation specifically requires AI companies to use systems that make it possible to identify AI-generated content.
Anthropic is not alone in this regulatory compliance effort. The company highlighted that other major model developers have signed the exact same Code of Practice and will be implementing their own watermarks moving forward.
Despite the industry-wide shift, user reactions have been intensely polarized. On Reddit, for example, one poster characterized this as a conspiracy against innocent Claude users, while another claimed, “The only reason you wouldn’t want this is to lie to people”.
The controversy has prompted highly visible pushback against the AI provider. Business Insider reports that “dozens” of users on X have claimed to cancel their Claude subscriptions as a result.