Wednesday, September 16

Programmers Discover Ways to Bypass Claude’s Text Watermarks

Emerging Challenges in AI Watermarking

Just four hours after Anthropic announced the integration of invisible watermarks into its Claude models, developer Guillaume Meyer had already devised a method to circumvent them. His code, designed to strip the watermarks from text generated by Claude, quickly gained traction on GitHub, amassing over 20,000 likes on X and attracting more than 100 contributors, many of whom are incorporating this technology into their own projects. An AI specialist commented on the rapid development, stating, “Anthropic is adding watermarks to its Claude texts… the issue is practically history just one day later,” alongside a visual of Meyer breaking free from chains while standing atop crumpled flags of the EU and Anthropic.

Following Anthropic’s announcement last week regarding the adoption of watermarks to comply with the European Union’s AI Act, Meyer and others began to delve into the workings of the watermarking system.

The Debate Over Watermarking

Some individuals attempt to bypass watermarks due to their disagreement with the principle that all AI-generated content should be labelled as such. Meyer explained that others, including himself, simply enjoy the technical challenge. Content writers and social media creators have also reached out to Meyer for assistance with the code.

The newly implemented regulations, which came into effect earlier this month, mandate that providers of models like Anthropic and OpenAI must label synthetic audio, images, videos, or texts so that machines can recognise the material as AI-generated; failure to comply could result in fines of up to 3% of annual revenue. Although the regulations prohibit providers from marketing evasion tools, there are no legal restrictions on independent tools.

READ:  Porter Nutrition: practical, evidence-based nutrition support in Manchester

“I am not against transparency; in fact, I advocate for content attribution. I merely believe that watermarks are a fundamentally flawed solution, presenting significant drawbacks and risks,” Meyer asserted. He is particularly concerned about the potential for false positives and the inability of watermarks to differentiate between light or heavy AI usage. As a native French speaker, he often utilises Claude and other AI tools like Grammarly for editing, raising the concern that using watermarks as evidence could lead to unjust outcomes. Even Anthropic acknowledges that its system can only estimate the likelihood of text being generated or modified by Claude. As Meyer points out, a company might reject a candidate or a researcher could face exaggerated accusations simply because a detector flagged their text.

How Watermarking Influences AI Output

Anthropic embeds the watermark invisibly, creating a pattern in Claude’s word and phrase selection that is undetectable to the human eye but can be identified by a machine programmed to seek it out. Some users are apprehensive that this may compromise the quality of Claude’s responses, though Anthropic maintains that it will not. The watermarking technique, known as SynthID, was developed by Google, which has been implementing it in its AI-generated content since 2023. Computer scientist Scott Aaronson proposed a similar approach while at OpenAI, but he claims the company opted not to implement it due to concerns that watermarks might deter customers from their product.

Innovative Approaches to Evasion

Meyer’s evasion method utilises a large language model (LLM) that does not apply watermarks, generating multiple rewrites by substituting words with synonyms and slightly rearranging the content. This approach relies on the availability of other large-scale language models that do not embed watermarks, which may not be reliable, as 190 organisations, including providers like OpenAI, Microsoft, and Meta, have signed the EU’s transparency code of practice. It remains to be seen how many of these labs will implement their watermarking strategies, which must be included in all new models launched from August onwards and integrated into existing models by December.

READ:  Economic Secretary to the Treasury's Address at UK Finance Event

While there is uncertainty surrounding the effectiveness of this tool until Anthropic releases the software it uses for watermark detection, understanding the basic approach behind SynthID-text, which informs Claude’s watermarking system, gives developers confidence in its functionality. Wayne Pan, Chief Technology Officer and co-founder of Silicon Valley-based AI startup Haimaker, incorporated Meyer’s open-source tool into his platform, sharing Meyer’s concerns about the implications of watermarked content, especially when only slight edits have been made.

A Growing Ecosystem of Evasion Tools

Other developers have also crafted their own evasion tools. Software engineer Erik Hughes managed to create a tool in just 15 minutes that utilises Claude to eliminate invisible characters and similar features, rearranging phrases within paragraphs and replacing several words with synonyms. Leon Chlon, a visiting researcher at the University of Oxford, suggests that one can remove watermarks by condensing Claude’s responses, translating them into a language such as Arabic—known for its significantly different semantics compared to English—and then translating them back. Anthropic itself has acknowledged that heavily edited, paraphrased, or translated content might not carry a watermark.

In a statement, an Anthropic spokesperson remarked, “We are adding watermarks to Claude’s outputs to comply with the EU AI Act, and other labs are taking similar steps.” They highlighted the challenges associated with identifying AI-generated text, asserting that this provides individuals with better tools for recognition.

“The text from compatible Claude models, including Claude Code output, will carry an invisible watermark, and this does not alter the meaning, quality, or readability of Claude’s responses,” the spokesperson added.

READ:  Meta Unveils Muse: An AI Assistant Designed to Complete Your Tasks While Safeguarding Your Data

Future Developments in Watermark Detection

Anthropic is currently exploring how to implement watermark detection in text and plans to launch a tool for this purpose; at that point, developers will finally be able to ascertain whether their methods are infallible. Additionally, the company continues to work on enhancing the watermarking system. Pan concludes, “They wanted to demonstrate that they are acting in good faith, but I do not believe it is possible to create a watermark that can withstand everything.”

Leave a Reply

Your email address will not be published. Required fields are marked *