Zero. That is the number of people who actually believe invisible watermarks in LLM output will stop the flood of AI-generated spam. Anthropic is essentially putting a “do not steal” sticker on a digital file and hoping the internet behaves. It is a gesture toward regulatory compliance, not a technical solution.
Invisible, machine-readable watermarks for generated text. C2PA provenance metadata for generated images. A pivot toward meeting European transparency mandates.
The driver here isn’t a sudden desire to save the internet from misinformation; it is the EU AI Act. Anthropic wants the European market, and they cannot afford to be the odd one out when the regulators start auditing provenance. According to The Verge, the company is pledging to mark text and images to ensure transparency. It is a classic corporate move: do the minimum required to satisfy the legal department while pretending it solves a systemic problem. This is the “Brussels Effect” in action, where a single region’s regulatory appetite forces a global change in product design, regardless of whether that change actually adds value to the end user.
On the technical side, text watermarking is notoriously fragile. Most of these systems rely on a technique called logit biasing, where the model partitions its vocabulary into “green” and “red” lists. By subtly favoring green tokens, the model creates a statistical fingerprint that a detection tool can spot, but a human cannot. But this signal is incredibly easy to kill. A user can just take the output, run it through a different, smaller model for a “rewrite,” or even spend two minutes manually swapping a few adjectives. It is a bit like trying to prove a photocopy is a copy by looking for a tiny ink smudge in the corner that the user can just crop out. Or maybe it is more like trying to hide a scent with cheap perfume; it might fool a casual observer, but it doesn’t change the underlying nature of the thing. Once the text is modified by even a small percentage, the “invisible” mark vanishes.
The image side is slightly different but equally futile. Anthropic is leaning on C2PA, a standard that attaches a digital manifest to the file. While this is a cleaner approach than pixel-level watermarking, it is fundamentally a voluntary label. Does anyone actually believe a bad actor—the kind of person actually deploying deepfakes or disinformation campaigns—is going to leave the metadata intact? A simple “save as” in a basic image editor or a quick pass through a scrubbing tool wipes the provenance clean. In fact, most social media platforms already strip this kind of metadata automatically to save space or protect privacy. The real-world friction here is almost zero. We are treating a metadata tag as if it were a physical seal on a vault, when in reality it is more like a “Made in China” tag on a t-shirt that can be snipped off with a pair of nail clippers in three seconds.
This move is more about brand management than actual safety (and they love their brand). By adopting these standards, Anthropic can maintain its image as the “responsible” lab while avoiding the heavy fines associated with the EU’s transparency rules. They are building a fence that looks great from the road but has a giant hole in the back. They want the prestige of being the “Constitutional AI” company without the actual difficulty of solving the attribution problem. By Q1 2025, we will see the first suite of “watermark scrubbing” tools specifically designed for Claude’s text output, likely packaged as “AI Humanizers” for students and lazy copywriters who want to bypass detection.
A corporate checkbox exercise.