Here is the uncomfortable thesis floating out of academic AI circles right now: the very tools being built to make language models safer are, structurally, a censor's toolkit. A 2026 position paper on arXiv by Sarah Ball and Phil Hackemann argues that modern alignment methods, designed to prevent harmful outputs, are converging on mechanisms that any authoritarian state would recognize as content control infrastructure. This lands the same week Joshua Kushner publicly warned Silicon Valley VCs that AI euphoria is warping judgment, and the same week Nature's Anthropic watermarking briefing asked whether provenance signals will actually change anything.
When Safety Infrastructure Becomes Suppression Infrastructure
The Ball and Hackemann paper does not accuse alignment researchers of bad faith. It accuses them of building systems whose architecture is indistinguishable from censorship at scale. Filters, refusal mechanisms, output classifiers: these are technically identical whether the goal is preventing a chatbot from writing bomb instructions or preventing it from discussing a government's human rights record. The paper is a position piece, not empirical, but the logic is airtight enough to make you uncomfortable. Meanwhile, Kushner's letter to Thrive investors frames the same period as one of historic capital misallocation driven by hype. The through-line: the industry is moving fast and encoding values it has not examined.
Watermarks, Alignment, and the Provenance Trap
Nature's briefing on Anthropic's new AI watermarking raises a parallel problem. Watermarks are sold as transparency tools, but they require the same classification infrastructure as content moderation. You cannot build a system that knows what AI generated without also building a system that can suppress what AI generates. A separate 2026 arXiv paper by Machidon et al. found that LLMs reach similar ethical conclusions to humans via completely different moral reasoning, meaning agreement is not alignment and surface compliance masks deep divergence. That divergence is exactly the gap through which a censor drives a truck. Kyle Raymond Fitzpatrick's essay on enshittification at Culture Slop made the same structural point about platform moderation years before alignment became the dominant frame: the tools built for safety always get captured by the entity with the most to lose from open speech.
The Discipline Problem That Kushner Cannot Name
Kushner's warning about investment discipline in AI is ultimately a warning about epistemic discipline. When a field is euphoric, it stops examining what it is actually building. The Ball and Hackemann paper is the academic equivalent of someone in the room raising their hand and asking if anyone has thought about the second-order effects. The Machidon et al. ethics paper adds empirical texture: models that say the right things are not models that reason the right way. That gap, between output and process, is where alignment becomes something else entirely.