Two papers landed this week on arXiv that are in direct, unresolved tension with each other. One proposes a new framework called Evaluative AI that would help humans make better decisions by having AI systems construct arguments rather than just produce outputs. The other argues that human oversight of AI becomes structurally impossible in high-stakes domains. Read together, they describe a system that wants to be trusted and cannot be supervised. The Vatican's art triennial is starting to look less eccentric by the hour.

Evaluative AI and the Argument for Argument

The paper by Yin, Miller, Potyka, Rago, and Toni on argumentative foundations for Evaluative AI proposes something elegant: instead of AI systems just producing recommendations, they should surface the reasoning structure behind those recommendations, making the logic auditable. This is distinct from explainability in the current XAI sense, which tends to produce post-hoc rationalizations. The EAI framework wants the argument to be the product, not the output with a justification stapled on. It is a serious proposal and it addresses a real problem: decisions made by opaque models are hard to contest.

When Oversight Becomes Structurally Untenable

The problem is that the Naito paper on content-judgment bypass in high-loss domains demonstrates that human-in-the-loop oversight fails precisely in the cases where the stakes are highest. The argument is structural, not contingent: when decisions need to be fast, when the domain is complex, and when errors are catastrophic, the conditions that make oversight meaningful are the same conditions that make it impossible. You cannot simultaneously require human review and require real-time response. These two papers are not describing different problems. They are describing the same problem from opposite ends, and neither resolves it. The GenAI newsroom study from Norwegian researchers, also released this week, found that journalism provides a useful case study of this catch-22: editors want to oversee AI-generated content but the production speed that makes generative AI valuable eliminates the time required for meaningful oversight. The governance problem is not a technical problem that better engineering will solve. It is a temporal problem baked into the logic of automation itself. Brewster Kahle's argument for public AI infrastructure, explored at Culture Slop, becomes more pressing in this context: if private systems cannot be supervised, the alternative may be public ones designed for transparency from the ground up.