Technology · Digital regulation
EU AI watermarking rules force Anthropic to alter word choices in generated text
Anthropic's compliance with the AI Act introduces invisible watermarks by substituting synonyms, prompting criticism that regulation treats language as interchangeable data rather than human expression.
Anthropic has quietly published the technical details of how it will watermark text produced by its Claude models, fulfilling a transparency obligation written into the EU AI Act. The method works by nudging the model's word selections along a statistical pattern governed by a cryptographic key known only to Anthropic. To the company's engineers, the substitutions are benign: 'grey' becomes 'overcast', 'veiled' becomes 'concealed', 'troop' becomes 'brigade'. To a writer, the changes are a quiet vandalism.
How the watermark works
The system does not insert visible markers or metadata tags. Instead it biases the probability distribution from which the model draws each token, creating a fingerprint that Anthropic can later detect by re-scoring the same text with its private key. The company argues that because the swapped words are near-synonyms, meaning is preserved. In a technical blog post accompanying the release, Anthropic researchers wrote that the alterations "do not change the semantic content" and that detection reliability improves with longer passages. They acknowledge, however, that certainty is probabilistic, not absolute.
The obligation originates in Article 50 of the AI Act, which requires providers of general-purpose AI models to ensure their outputs are "marked in a machine-readable format and detectable as artificially generated or manipulated." The regulation entered into force in August 2024 and its transparency provisions apply from August 2025. Anthropic's implementation is the first detailed public example of how a major model provider intends to comply.
A writer's objection
Jeff Jarvis, author of the forthcoming book Hot Type and a professor at the Craig Newmark Graduate School of Journalism, has made the most sustained public critique of the approach. His argument is not that watermarking is technically flawed, but that its philosophical premise is corrosive. By treating 'grey' and 'overcast' as interchangeable, he says, Anthropic and the legislators who mandated the measure declare language itself to be fungible data.
"I try to select my words as carefully as I can for a number of considerations: style, tone, rhythm, avoiding repetition, but most of all meaning," Jarvis writes. "You may quibble with my choices, but they are mine, not yours; that's what makes my writing mine, to communicate what I wish to communicate." The watermark, he argues, severs the link between author and intention, replacing deliberate selection with algorithmic permutation.
Demonstrating the loss
To make the point concrete, Jarvis fed the opening paragraph of Hot Type into Claude with a request for a synonym-substituted version that preserved structure and meaning. The original describes the Linotype machine's exposed innards, "its troop of idiosyncratically shaped cams, gears, belts, arms, and wheels marching to instructions foreseen more than a century before by their inventor, German-born watchmaker Ottmar Mergenthaler." Claude returned: "its brigade of distinctively shaped cams, gears, belts, arms, and wheels progressing to directives anticipated more than a century earlier by their creator, German-born horologist Ottmar Mergenthaler."
The substitutions are defensible in a thesaurus. 'Troop' to 'brigade', 'idiosyncratically' to 'distinctively', 'marching' to 'progressing', 'foreseen' to 'anticipated', 'watchmaker' to 'horologist'. Yet the texture evaporates. The military metaphor of 'marching' disappears. The anachronistic charm of 'watchmaker' yields to the clinical 'horologist'. The rhythm stretches. Jarvis calls the result "fine" in the way a replicated chair is fine, functional, but absent the grain.
Regulation as moral panic
Jarvis traces the watermarking mandate to what he calls a "prejudice against the technology: fear unto moral panic." The AI Act was negotiated amid intense public anxiety about deepfakes, disinformation and academic cheating. Lawmakers in the European Parliament and the Council pushed for strong transparency guardrails, arguing that citizens have a right to know when they are reading machine output. The final text reflects that urgency: general-purpose model providers must implement "effective, interoperable, non-discriminatory and reliable" marking techniques.
The irony, in Jarvis's view, is that a measure designed to protect the sanctity of human text achieves the opposite. By normalising the idea that words are swappable tokens, it accelerates the commodification of writing that large language models have already set in motion. "At the very moment when LLMs commodify writing and writers, this is just another kick to the kidneys for a challenged craft," he writes.
Historical perspective on writing technology
Hot Type itself is a history of how writing tools shape thought. Jarvis cites Friedrich Kittler's observation that the typewriter altered how Nietzsche, Hemingway and Hesse composed. He recounts his own trajectory from typewriter to computer to PostScript, noting that each technology changed his methods without making his prose less human. The argument is not that AI should be unregulated, but that the current regulatory impulse mistakes a cultural transition for a forensic problem.
"We should debate what the impact of this next technology, AI, might be on creativity and perception, but we should give that process time," Jarvis writes. "Rushing to regulatory and technological solutions to a problem not yet fully understood is foolish." The AI Act's timetable, proposed in 2021, agreed in 2023, applied from 2025, reflects political momentum more than deliberative patience.
What the watermark cannot solve
Technical limitations compound the cultural critique. Watermarking only works on text generated directly by a compliant model. It does not survive substantial human editing, translation through a non-compliant system, or paraphrase by a second model. Adversarial actors can strip the signal by rewriting. Researchers at ETH Zurich and elsewhere have shown that detection rates drop sharply under even light paraphrase attacks. The European Commission's AI Office is tasked with monitoring compliance, but the regulation does not specify a minimum detection threshold.
Moreover, the watermark applies only to providers who choose to comply within EU jurisdiction. Open-weight models released elsewhere can be fine-tuned without the marking layer. The result is a fragmented landscape where regulated outputs carry a statistical fingerprint while unregulated equivalents flow freely, a dynamic familiar from copyright enforcement and data protection regimes.
Industry response and next steps
Other major model providers have not yet published comparable watermarking designs. OpenAI has signaled support for provenance standards such as C2PA but has not detailed a token-level watermark for ChatGPT. Google's SynthID marks images and audio; its text watermark remains in research. The AI Act's implementing acts, due to be drafted by the Commission with advice from the European AI Board, may eventually prescribe technical standards. Until then, Anthropic's approach is the de facto reference implementation.
Sources
People mentioned
Jeff Jarvis
Organisations
Anthropic · European Commission · European Parliament