Anthropic's Claude AI Text Watermark: A New Feature for Transparency

·

A new feature has been implemented in the future versions of Anthropic’s Claude language model, designed to help identify whether a given piece of writing was likely produced by its AI. This change is meant to comply with EU regulations and provide transparency into the use of artificial intelligence-generated content. The company has explained how this watermark works and why it’s necessary.

The feature exploits the countless small decisions made by the language model as it generates text, rather than using arbitrary random numbers. Instead, the watermarked version of Claude bases these decisions on a cryptographic key combined with the preceding text. This results in a subtle statistical pattern spread across the response that is invisible to human readers but detectable to anyone with the matching key.

The watermark allows users to estimate the probability that Claude generated the text, providing a way to identify AI-generated content. However, it’s essential to note that this feature doesn’t change the normal writing output of Claude and is meant to be indistinguishable from unwatermarked responses.

Anthropic has stated that the implementation carries no cost to output quality, with internal testing showing no measurable difference in creativity, accuracy, or readability between watermarked and unwatermarked responses. The company also pointed out that Google DeepMind’s original research on the underlying technique supports these findings. Additionally, Anthropic noted that its research showed no statistically significant shift in user satisfaction when a similar watermark was tested on live traffic.

The feature adds no extra tokens, meaning it doesn’t slow Claude down or make it more expensive to use. However, there are limitations to the watermark’s effectiveness. It only works when a model is choosing among several equally valid options, so text with little room for variation carries a much weaker signal. Detection also grows less reliable on very short passages since there’s simply less pattern to analyze.

For instance, in cases where Claude has limited choices or produces hard factual statements, precise code, or math answers, the watermark may not be effective. Similarly, lightly edited or proofread human writing may carry little to no detectable trace of the watermark, as most of the original wording remains unchanged. A sufficiently heavy rewrite can even remove the watermark entirely.

Anthropic emphasized that the watermark cannot be traced back to a specific user, account, or conversation and does not establish authorship or ownership over content. All it can tell you is the likelihood that Claude was involved in producing or editing it at some point. The company also distinguished its approach from third-party AI-detection tools, which typically rely on spotting stylistic patterns in AI writing rather than checking for an embedded signal tied to a private key.

The rollout of this feature is tied to regulation rather than a purely voluntary move. Anthropic signed the European Union’s Code of Practice on Transparency of AI-Generated Content in July 2026 and plans to extend it to older Claude models over the coming months. The company also stated that it will soon offer a separate API allowing anyone to check whether a piece of text carries Claude’s watermark.

Anthropic is implementing this feature globally, as they don’t yet have a reliable way to apply the watermark only within the EU. This decision was made in response to an EU AI Act requirement, effective August 2, that AI providers mark generated text. The company aims to provide transparency into its use of artificial intelligence-generated content and comply with regulatory requirements.