Two stories about who shapes AI landed this week, and together they outline a structural problem that neither the tech industry nor the legal system has fully reckoned with. A new arXiv paper by Elena Kopteva and Vitaliy Hlynianyi-Zhuk identifies what they call rater state bias in RLHF preference data: the geographic and cultural location of human raters systematically skews what AI systems learn to reward. Separately, a U.S. federal judge approved Anthropic's $1.5 billion settlement with authors whose books were used without permission to train AI models. Two different bias problems. Same underlying question: whose knowledge, and whose judgment, is being baked in?
The Rater as Ghost in the Machine
The Kopteva and Hlynianyi-Zhuk paper proposes an audit framework for detecting rater state bias, which is a polite way of saying that right now, most RLHF pipelines have no systematic way to know whether their preference data reflects a genuinely diverse set of human values or just the preferences of whichever contractor pool was cheapest and most available. This is not a new critique, but the formal audit framing matters. A 2023 paper in Proceedings of the ACM on Human-Computer Interaction by Bender et al. argued that the values embedded in language models are not neutral encodings but political acts. The rater state bias paper gives that argument a measurable surface.
The Author Settlement's Hidden Logic
The Anthropic settlement is the legal system catching up to the training data problem from a different angle. Authors did not consent to their prose becoming the substrate for a model's stylistic preferences, which are then used to evaluate other prose in RLHF pipelines. The circle is closed and the original sources are invisible. Brewster Kahle's conversation on public AI and the global brain argued that the architecture of knowledge access is a political choice. Whose writing trains the model that trains the model? That question now has a $1.5 billion partial answer, but the rater bias paper reminds us the problem runs deeper than copyright. It runs through every preference signal collected from every annotator in every location, shaping every response the model will ever give.