Anthropic Introduces Watermarks in Response to EU AI Act

Evasion Possible by Rephrasing Words and Sentences

Emerging Removal Tools Highlight Limits of Perfect Detection

The movement to add watermarks to content created by generative artificial intelligence (AI) is spreading across AI companies. The purpose is to establish mechanisms for identifying AI-generated products and enhance transparency, but there are concerns that the simultaneous development of methods and technologies to bypass such detections limits their effectiveness as foolproof identification tools.

Anthropic. Photo by AFP Yonhap News

Anthropic. Photo by AFP Yonhap News

View original image

According to industry sources on August 13, Anthropic began applying machine-readable watermarks to its Claude models to comply with the transparency requirements of the European Union (EU) AI Act, which took effect on August 2. This approach embeds markers within the text itself—imperceptible to humans but detectable using specialized tools. Anthropic plans to sequentially apply watermarks to all of its AI models in the future and will release detection tools and technical documentation that third parties can use to verify them.


The adoption of watermarks is becoming an industry-wide trend. Google applies watermarks to text and images generated by Gemini using "SynthID." OpenAI is embedding content provenance information into images using "C2PA" metadata. AI music service Suno and newsletter service Substack are also implementing ways to identify AI-generated content.


However, industry experts believe such efforts by AI companies are unlikely to be completely effective, as it is possible to evade detection simply by changing words or restructuring sentences. The watermarking approach chosen by Anthropic involves subtly increasing the probabilities of certain words being selected by the AI, creating statistical patterns throughout the text. Detection tools then analyze the frequency of specific words to determine the presence of a watermark.

Watermarks Spread Across AI Content... Are They Easily Removed? View original image

Alex Zhu, Chief Technology Officer (CTO) of AI detection firm GPTZero, pointed out, "If someone performs intensive paraphrasing that rewrites both words and sentence structures, the watermark is unlikely to survive." Google also stated at the release of SynthID that detection performance may be lower for short texts, rewritten or translated content, and answers to factual questions.


There are also concerns that watermarks could compromise the quality of generated material. CTO Alex Zhu noted that artificially adjusting the probability of words originally chosen by the AI could restrict word selection and potentially lower text quality.


Technologies that neutralize watermarks are also advancing. A prime example is "watermark stealing," which involves reverse engineering the watermarking rules. Researchers at the Swiss Federal Institute of Technology Zurich conducted experiments in which they repeatedly queried AI models with embedded watermarks to infer these rules. As a result, they recorded an average success rate of 80% in attacks that either erased existing watermarks or injected fake watermarks into human-written text, all at a cost of less than 50 dollars.



Tools that ordinary users can access for watermark removal are also emerging. In June, an open-source tool capable of stripping AI-generated image indicators was released on the development platform GitHub. This technology can remove invisible watermarks such as Google’s SynthID, C2PA metadata, and even watermark logos.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing