Cybersecurity

Tools for Removing Watermarks from AI-Generated Text Spread Without Evidence of Effectiveness

Within days of Anthropic announcing the activation of invisible watermarks in text produced by Claude, a group of open-source projects and commercial services emerged claiming to bypass these watermarks. However, these claims cannot currently be verified, as Anthropic has not published the watermarking mechanism or a tool to detect it.

2026-08-13
5 min read
8 views
فريق تحرير certi.news
Tools for Removing Watermarks from AI-Generated Text Spread Without Evidence of Effectiveness

The internet has seen the emergence of a rapidly growing market for tools that claim to remove watermarks from text produced by artificial intelligence, days after Anthropic announced that it had activated invisible watermarks in everything Claude writes. This market includes an open-source project with more than 4,500 stars on GitHub, a group of recently registered websites, and commercial services specializing in bypassing AI-generated-content detection tools.

However, it is currently impossible to verify these tools’ claims regarding the removal of text watermarks, because Anthropic has not yet published details of how its mechanism works, nor has it launched a public detection tool that could be used to determine whether cleaned text still carries the watermark.

Multiple Projects and Services

Among the most prominent projects is the watermarks-remover tool, licensed under the MIT license and developed by Guillaume Meyer, founder of Memo. The tool began as a skill designed for a Claude agent, then announced support for Claude, Gemini, and SynthID-Text, in addition to OpenAI’s provenance surfaces and open-weight models that use Kirchenbauer-style watermarks.

Other repositories exist alongside it, such as claude-watermark-cleaner, remove-ai-watermarks, and noai-watermark, as well as recently launched websites including claudewatermark.com, claudewatermark.rip, gptcleanup.com, and claudewatermarkremover.app. StealthGPT, a service that sells tools for bypassing AI-generated-content detection, has also added a Claude watermark removal tool to its service pages. Human Writes likewise advertises bypassing Turnitin and GPTZero in essays and assignments, while claiming to remove the Claude watermark, despite including a notice urging users to use the service in accordance with academic-integrity policies.

What Do These Tools Actually Remove?

The tools handle three different types of data. Removing hidden characters, such as zero-width characters, bidirectional control characters, special Unicode characters, and spaces that resemble ordinary spaces, is a process that can be verified and measured. The same applies to removing C2PA, EXIF, and XMP metadata from PNG, JPEG, SVG, PDF, DOCX, ODT, HTML, and Markdown files. However, this data usually does not survive after a file is resaved, its format is converted, or a screenshot is taken.

The text watermark itself, however, does not exist in hidden characters; rather, it is linked to the words selected by the model. Therefore, the known way to deal with it is to substantially rewrite the text using another model. Meyer explains that his tool currently removes only metadata, and that removing the actual watermarks may come later, but is not currently available.

Some commercial websites promise to produce clean, undetectable text, while some results are measured against general detection tools rather than Anthropic’s watermark, for which no public detection tool exists. After cloning the projects and reading the code, Pasquale Pillitteri also found that one popular text-cleaning tool left unchanged a common technique for carrying hidden data, allowing the payload to be decoded and recovered after the tool claimed to have cleaned the text.

The Reason for Watermarking and Supply-Chain Risks

Anthropic said that text outputs from models released on or after August 2, 2026, carry an imperceptible watermark embedded in the phrasing, while supported file types receive signed C2PA metadata. The labeling is applied at the model level through the API, claude.ai, Claude Code, Claude Cowork, and Claude Tag, as well as through AWS, Google Cloud, and Microsoft Foundry.

This measure is linked to Article 50 of the European Union’s AI Act, whose application took effect on August 2, with penalties that may reach €15 million or 3% of global revenue. Anthropic explains that detecting the watermark means the content was processed by Claude, not necessarily that it was written entirely by it; the watermark may appear when the model is used for proofreading, translation, or summarization. It also noted that extensive editing, rewriting, and translation may remove the watermark, and that it will support third-party detection and publish technical documentation later.

These developments also highlight supply-chain risks. The watermarks-remover tool is installed as a skill for a software agent and can be configured to clone a third-party research repository and download a file approximately 220 megabytes in size. BleepingComputer did not test or audit the tools mentioned, so they should be treated as untrusted code from the internet, particularly given the absence of an independent means of verifying their promises.

News source
BleepingComputer
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news