In an era where artificial intelligence-generated content is becoming increasingly prevalent, identifying the origin of such material is crucial. To combat potential misinformation and enhance transparency, Anthropic has introduced invisible watermarks for AI-generated text through its Claude models. This innovative step aims to give consumers the tools necessary to distinguish between human-written and AI-generated content easily.
Recently, you might have noticed watermarks appearing on AI-generated imagery, indicating their artificial origins. Following suit, AI-generated text is now being watermarked, allowing for easier recognition of online content produced by AI.
Anthropic, the organization behind Claude, has recently committed to the EU’s Code of Practice on Transparency of AI-Generated Content and elaborated on how this watermarking process functions. The watermarks embedded in the text are designed to be machine-readable while remaining unobtrusive for human readers. “When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response,” explains the company.
This subtle watermark is resistant to copying and pasting and can even endure minor edits, making it challenging to bypass. However, the watermark may not persist if the text is overly brief or subjected to extensive modifications.
Every Claude model launched in the EU on or after August 2, when the Code took effect, automatically incorporates this watermarking. The technology operates within the model, making it applicable across various platforms, including Claude API, Claude Code, Cowork, and Tag, regardless of whether the service provider is AWS, Google Cloud, or Microsoft Foundry.
Moreover, Anthropic is working to extend watermarking functionality to older models predating August 2, with plans for a global rollout. The company is also developing tools for detecting these watermarks, which will be available shortly.
It is important to note that the watermark will be applied to all text generated by Claude models, even when the AI is not the primary author. For instance, when using Claude for tasks like translation, summarization, formatting, or proofreading, the output will still carry the watermark. This application goes beyond the stipulations of the Code of Practice, which states an exception for systems performing assistive functions that do not significantly alter the input data. Anthropic has decided to apply watermarks universally as a precaution.
In addition to text, AI-generated images will also feature watermarks that, unlike past implementations, are difficult to remove. These images will be accompanied by provenance metadata in formats like JPG, PNG, and SVG, adhering to open standards established by the Coalition for Content Provenance and Authenticity (C2PA).
While enhancing detection of AI-generated content benefits consumers significantly, it also serves AI companies, especially in a digital landscape increasingly crowded with bot-generated comments and articles. As organizations may begin to train new AIs on text produced by older models, this initiative ultimately leans towards a positive outcome for consumers.