cs.LGMay 13, 2026

The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next

Authors: Adil Amin

Organizations: ZEHEN Labs

Abstract

Leaderboards rank frontier models on independent axes but do not reveal whether capabilities reinforce or trade off across releases -- and at the frontier, this interaction is the more informative signal. We decompose paired SWE-bench and GPQA Diamond scores into a population coupling trend and per-release residual (hh-field) that diagnoses capability emphasis from two public benchmark scores. Across 34 models from 10 labs (2024--2026), capabilities cooperate (r=+0.72r = +0.72, p<106p < 10^{-6}), but cooperation varies systematically: per-lab coupling slopes span 5×5\times (Google 1.151.15 vs. DeepSeek 0.230.23), and labs pivot -- DeepSeek reversed from reasoning-rich to coding-first (Δh=15.9Δh = 15.9~pp); Anthropic oscillates between coding excursions and recovery. The population regression serves as an isocline phase boundary: the same (a/b)B1\sqrt{(a/b)\cdot B_1} classifier that identifies the base-scale coupling transition [Amin, 2026] classifies frontier models and already detects mixed-phase behavior at the next transition (two models below the GPQA--IFEval isocline). The hh-field is not just diagnostic -- it tells you what to change. Pretraining establishes coupling at 0.8710.871 while RLHF adds 0.0810.081 [Amin, 2026]: pretraining-level shifts are permanent (DeepSeek's four-release reversal persists), post-training shifts are reversible (Anthropic's three coding excursions each recover within one release), and inference compute alone shifts hh by +7.8+7.8~pp without retraining. Knowing which component dominates determines whether to retrain or wait. We provide a three-step diagnostic (locate, classify, predict), a per-lab measurement-priority table, and seven falsifiable predictions with timestamped criteria. Five post-cutoff releases fall within the 95% prediction interval. Code, data, and an interactive dashboard: https://zehenlabs.com/cape/.

Explore similar work

CardsList
  1. You Don't Need to Run Every Eval

    Jun 22, 2026Yuchen Zeng, Dimitris PapailiopoulosModel EvaluationScoring

  2. The Capability Frontier: Benchmarks Miss 82% of Model Performance

    Jun 25, 2026Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8FrontiersPareto Frontier