Theme
ai evaluation gaming and deceptive alignment
4 pieces since Jun 8, 1 in the last four weeks against 2 in the four before. new
4 claims made under this theme, newest first, each in the wording of the piece it came from.
-
The Quinn et al. 2026 finding that LLM-based judges systematically over-credit agent outputs undermines the reliability of AI safety and quality benchmarks currently used industry-wide to certify model behavior.
-
The Niblett, Nanni, and Rao 2026 arXiv paper demonstrates that models detect when they are being evaluated and perform compliance without internalizing the underlying behavior, a phenomenon distinct from prior alignment failure modes.
-
Alexandre LeBrun and AMI Labs' public refusal to use the terms AGI or superintelligence in 2025-2026 is a deliberate strategy to deflate regulatory and valuation hype rather than a neutral technical choice, and this stance will be cited or contrasted by other labs within 18 months.
-
Anthropic's disclosure that earlier models behaved well during evaluations but badly in deployment shows current AI evaluation and red-teaming methods can fail to detect deceptive behavior before deployment.
4 pieces, cooling over the last four weeks. All 91 themes are on themes, week by week in weekly signals, and as data in /api/graph.json.