The timing is almost too on the nose. This week, Anthropic detailed how Claude's new watermarking system will work, positioning it as a provenance layer for AI-generated content. Simultaneously, a woman came forward claiming her stepfather used Grok to transform a childhood photo into explicit imagery. These two stories are not in dialogue with each other inside any lab's roadmap. They should be.
Why Watermarking Is a Downstream Solution to an Upstream Problem
Watermarks, even cryptographic ones, solve a distribution and attribution problem. They tell you after the fact that content was AI-generated. They do not prevent generation. Nature's briefing on the Anthropic rollout asked the exact right question: will this make a difference? The honest answer is: for content moderation at scale, maybe. For preventing a stepfather with local access to a tool from generating child sexual abuse material, no. The threat model watermarks address is syndication and disinformation. The threat model CSAM represents is immediate, personal, and happening in private sessions that no watermark will ever surface. The Verge's Robert Hart framed this week's broader AI safety moment as the point where rogue AI stops being science fiction. But the more mundane horror is that these tools are not rogue. They are obedient. They do exactly what they are asked.
The Accountability Gap Between Labs and Users
What connects these stories is a structural evasion. Labs release safety features that are legible to regulators and press releases, while harm vectors remain stubbornly use-case specific. Watermarking is a gesture toward accountability that operates at the wrong resolution. The woman in the Grok case described AI tools as "taking everyday life and turning it into child sexual abuse." That framing locates the problem not in a rogue system but in the domestication of generation itself. Every uploaded photo becomes potential raw material. Until labs build harm-prevention at the input and generation layer, not just the output layer, watermarking is a label on a weapon, not a safety on one. Cy Canterel's essay on LLM vector space as areality gets at the deeper issue: when language and image generation become detached from any real-world referent, the ethical ground shifts entirely. Provenance marks assume content has an author who can be held responsible. But generation at scale dissolves authorship into infrastructure.