Anthropic Clarifies How Claude's Watermarks Will Work

·

A recent blog post from Anthropic aims to address concerns about its plan to watermark text generated by the chatbot Claude. The move is in compliance with the EU AI Act’s Transparency Code, which requires AI companies to use systems that make it possible to identify AI-generated content.

Claude users have been debating this decision on Reddit and other platforms, with some expressing concern that watermarks could be seen as a ‘conspiracy against innocent Claude users.’ Others argue that having transparent information about AI-generated text is essential for maintaining trust in these tools.

Anthropic’s new post explains the watermarking concept by comparing it to making low-stakes choices. For example, when describing the weather, Claude might choose between ‘overcast’ and ‘grey’. In this case, the model can create a pattern that is undetectable to readers but detectable with a key.

The company emphasizes that watermarks do not impact the quality of Claude’s output. To a reader, a watermarked response should be indistinguishable from an unwatermarked one.

Anthropic will use the SynthID-Text approach developed by Google DeepMind in 2024 to implement its watermarking system. Additionally, it plans to release a watermark detection API for developers and researchers.

The company also clarifies that watermarking is distinct from AI detection approaches offered by companies like Pangram. These methods look for ‘tells’ in the writing, such as specific constructions or word choices, whereas watermarks are unique identifiers embedded within the text itself.

Some users have expressed concerns about whether it’s possible to rewrite or edit a watermarked text to remove the watermark completely. Anthropic acknowledges that this might be feasible with significant editing efforts but notes that ‘a complete rewrite where every word is replaced’ would likely eliminate the watermark entirely.

The company also addresses the issue of code, stating that it should have less of a watermark than other text because the model needs to create working code and has limited freedom in making choices. However, Anthropic suggests that watermarks can still be used in certain areas, such as comments within code, but with negligible effect on actual code produced.

Anthropic’s move is part of its compliance efforts under the EU AI Act’s Transparency Code. The company notes that other major model developers have signed this Code and will implement their own watermarks to ensure transparency in AI-generated content.