Someone at Meta built an internal leaderboard called Claudeonomics, ranked the company's top 250 AI token consumers, and handed out titles. Rita McGrath at Fast Company is not surprised the results were underwhelming. She should not be. Measuring token consumption to proxy for AI-driven productivity is like measuring page views to proxy for journalism quality. The metric is real. The thing it is supposed to represent is not what it is measuring.
Why Performance Metrics Fail AI Systems
A 2026 paper in arXiv CS.AI by Du et al. on generation-aware evaluation as actionable feedback argues that current LLM evaluation methods are coarse-grained and decoupled from what the model actually produces: they measure outputs in isolation from the generative context that produced them. This is the same problem at organizational scale. Claudeonomics measures consumption, not output quality. It rewards the person who prompts most, not the person who prompts best. A 2026 arXiv paper by Rai et al. on second-order social reasoning in LLMs found that most AI alignment work focuses on first-order norms, what the model does, rather than second-order norms, what the model understands about what is expected of it socially. The same gap exists in corporate AI adoption: first-order metrics of usage, zero second-order metrics of judgment.
The Gamification of Intelligence
The deeper problem with Claudeonomics is that it turns intelligence augmentation into a leaderboard, which immediately changes the behavior it is trying to measure. People optimize for the metric, not for the work. This is Goodhart's Law applied to cognitive labor. AI-designed drug trials are this week's counterexample: here, AI is doing genuinely measurable work, shortening timelines, identifying candidates a human researcher would miss. The difference is that the output is testable against reality. Most corporate AI deployments have no equivalent reality test. They have Claudeonomics. That is the crisis the underwhelming metrics are telling you about.