Skip to content

Europe

Independent · Brussels & Berlin

Technology · Digital regulation

EU AI Act labelling rules take effect but researchers warn of boomerang effect

Article 50 transparency obligations became enforceable on 2 August, requiring visible icons and machine-readable watermarks on synthetic media, yet evidence suggests labels alone may increase trust in unmarked deepfakes.

By , Technology Editor

Published

7 min read

The European Union's most ambitious attempt yet to curb synthetic media deception became enforceable on 2 August, when Article 50 of the AI Act moved from legislative text to legal obligation. From that date, any company providing or deploying AI systems that generate or manipulate images, audio, video or text resembling real people, places or events must disclose the synthetic origin of that content. The rules apply equally to firms headquartered outside the bloc if their output reaches users inside it.

What the law now requires

The transparency obligations are specific. Visual deepfakes must carry a clearly visible icon at the point of first encounter, embedded so the mark survives sharing and downloading. Audio deepfakes need an audible disclaimer at the start. Text on matters of public interest that has not undergone human review or editorial control must also be flagged. The European Commission has published a set of standard icons, though their use is optional provided the disclosure requirement is met in another equally clear manner.

Beyond visible labels, the Act demands invisible, machine-readable watermarks on any content generated or even processed by an AI system. This technical requirement is designed to allow automated detection downstream, by platforms, fact-checkers or archival tools. Anthropic, the US-based AI laboratory, has already stated that all Claude chatbot models launched in the EU after 2 August will support such marking, and that it is working to retrofit earlier models.

Scope, penalties and the personal-use gap

The territorial reach is broad. A startup in Toronto or Singapore that serves synthetic media to a user in Berlin falls under the same regime as a Paris-based developer. Non-compliance carries a maximum fine of €15 million or 3% of total worldwide annual turnover, whichever is higher. That ceiling mirrors the sanction structure of the General Data Protection Regulation and signals the seriousness with which the Commission treats the new transparency layer.

A notable exemption weakens the net. The Act does not apply to individuals acting in a personal or non-professional capacity. Someone generating a deepfake video on a home computer for private amusement, or sharing it in a closed messaging group, faces no legal obligation to label it. Malicious actors operating anonymously or pseudonymously are similarly unlikely to self-disclose. The regulation therefore targets the supply side, model providers and commercial deployers, rather than the demand side of deception.

The boomerang effect: when labels backfire

The central policy assumption is that visible disclosure reduces harm by alerting audiences. A growing body of behavioural research challenges that assumption. Studies cited by the Commission's own advisers suggest a 'boomerang effect': repeated exposure to labelled synthetic content conditions viewers to treat the absence of a label as a proxy for authenticity. In experiments, participants shown a mix of labelled and unlabelled deepfakes became more likely to trust the unlabelled ones, even when those were also synthetic.

The mechanism is straightforward. Human cognition relies on heuristics. If a platform consistently flags AI output with a badge, the badge becomes a signal: 'this is the AI stuff'. Content without the badge is implicitly categorised as 'not AI stuff'. Malicious actors who strip watermarks, evade labelling APIs or simply use open-source models that do not implement the marking standard gain a credibility advantage precisely because their output lacks the warning.

Social cues override technical signals

Labels also compete with stronger social signals. Research shows that virality metrics, view counts, share numbers, comment threads endorsing a narrative, can override explicit authenticity warnings. A deepfake video bearing an EU-mandated icon but accompanied by thousands of supportive comments and a high view count is often judged more credible than a genuine video with low engagement. The label becomes background noise; the crowd becomes the credential.

This dynamic is not new. Fact-checking labels on social platforms have exhibited similar limitations. The difference is that the AI Act makes labelling a legal baseline for commercial providers, potentially creating a false sense of security among regulators and the public that the problem is 'being handled'.

Technical arms race and the open-source blind spot

Machine-readable watermarks are touted as the technical backbone of enforcement. Yet the ecosystem of open-weight models, Llama derivatives, Stable Diffusion forks, community fine-tunes, operates largely outside the commercial API channels the Act can easily police. A developer who downloads a model weights file and runs inference locally produces content with no embedded watermark unless they voluntarily add one. The Commission's implementing acts will need to address whether distributing such models constitutes 'providing an AI system' under the Act, a question that remains unsettled.

Anthropic's commitment to watermark Claude output is a compliance signal from a well-resourced proprietary vendor. Smaller firms, academic labs and hobbyist communities have fewer incentives and less capacity to build robust marking pipelines. The result may be a two-tier synthetic media landscape: watermarked output from regulated commercial services, and unmarked output from the long tail of open-source deployment.

Artistic and satirical carve-outs

The Act acknowledges that not all synthetic media is deceptive. Content that is 'evidently artistic, creative, satirical or fictional' need only be disclosed in a manner that does not disrupt the presentation or enjoyment of the work. This carve-out is deliberately vague. A satirical deepfake of a politician shared on a comedy site may qualify; the same clip clipped and reposted on a partisan channel without context may not. The boundary will be litigated, and platforms will face pressure to over-label to avoid liability.

Literacy, not labels, as the long-term defence

The consensus among researchers is that labelling is a necessary but insufficient layer. As generative models improve, the perceptual cues that once betrayed synthetic media, asymmetrical earlobes, temporal flicker, phoneme mismatches, are disappearing. The next generation of diffusion and transformer models produces output that is statistically indistinguishable from camera-captured reality at the pixel and waveform level.

That shift moves the detection problem from perception to provenance. Who created or shared the content? Is the source credible? Does the claim make sense in context? Can it be verified independently? These are media literacy questions, not technical ones. The Commission has funded AI literacy initiatives under the Digital Europe Programme, but funding remains a fraction of what is spent on model development. Without sustained investment in public education, the labelling regime risks becoming a regulatory ritual that satisfies compliance officers while leaving citizens no better equipped.

Sources

  1. The Conversation

    theconversation.com · 2026-08-16

Organisations

European Commission · Anthropic

Related analysis

Selected because they share topics with this article

The newsletter

One important European story. Explained properly.

Delivered to your inbox on the days we publish. No daily digest, no push notifications, no advertising.

We store your address only to send the briefing. Unsubscribe in one click.