SciExam for ENSO: Can AI Agents Build Climate Models?
Organizations: The Ohio State University · Yale University · University of California, San Diego · Stanford University · Princeton University
Abstract
Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO's warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under varied information suggest that its scores do not come from recalling the dated observational record and that the information it receives shapes how it builds its model. SciExam for ENSO can thus evaluate agent research where no answer is known, and the results suggest that agents can already build competitive models whose structures bear on questions that scientists still debate.
Figures & tables
| Benchmark | Long horizon | Physical system | No known answer | Hidden scores | Scientific evaluation |
| Long-horizon agent tasks | |||||
| Agents’ Last Exam [ 27 ] | ✓ | ✗ | ✗ | ✓ | ✗ |
| MLE-bench [ 6 ] | ✓ | ✗ | ✗ | ✓ | ✗ |
| RE-Bench [ 32 ] | ✓ | ✗ | ✓ | ✗ | |
| ALE-Bench [ 11 ] | ✓ | ✗ | ✓ | ✗ | ✗ |
| EdgeBench [ 36 ] | ✓ | ✗ | – | ||
| Dimension | What it tests | Measure | Panels |
| Statistical properties | Whether a long free run of the model behaves like the observed ENSO | Probability density functions (PDFs) | a, e |
| Seasonal variance | b, f | ||
| Autocorrelation functions (ACFs) | c, g | ||
| Frequency of different event types | d | ||
| Event strength against location | h | ||
| Dynamical consistency | Whether the variables are coupled as observed | , , reconstructed from , | i–k |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Final checkpoint | Best checkpoint |
| Claude Fable 5.1 | 0.515 | 0.531 |
| Kimi K3 | 0.479 | 0.500 |
| GPT-5.6-Sol | 0.443 | 0.466 |
| GPT-6 Astra | 0.442 | 0.475 |
| Claude Opus 5 | 0.440 | 0.508 |
| GPT-5.5 | 0.403 | 0.404 |
| System | Scaffold | Route |
| Claude Fable 5.1 | Claude Code 2.1.270 | Subscription |
| Kimi K3 | Claude Code 2.1.223 | OpenRouter (Moonshot AI) |
| GPT-5.6-Sol | Codex 0.146.1 | Subscription |
| GPT-6 Astra | Codex 0.154.0 | Subscription |
| Claude Opus 5 | Claude Code 2.1.223 | Subscription |
| GPT-5.5 | Codex 0.146.1 | Subscription |
| System | Route | Calls | Prompt tokens | Cached | Cost (USD) |
| Kimi K3 | OpenRouter | 954 | 120.3 M | 99.2% | 43.69 |
| GLM-5.2 | OpenRouter | 669 | 65.4 M | 98.0% | 21.48 |
| Qwen3-Max | OpenRouter | 3,250 | 308.1 M | 95.3% | 132.37 |
| DeepSeek V4 Pro | OpenRouter | 589 | 48.7 M | 93.2% | 7.67 |
| MiniMax-M3 | OpenRouter | 1,011 | 83.1 M | 98.2% | 5.97 |
| DeepSeek V4.1 Flash | OpenRouter | 752 | 62.0 M | 98.4% | 2.43 |
| Symbol | Value | Symbol | Value |
|---|---|---|---|
| Symbol | Definition or value |
|---|---|
| Symbol | Value | Symbol | Value |
|---|---|---|---|
| Symbol | Definition or value |
|---|---|
| Composite | Best checkpoint | |||||
| Condition | Access | Best | Final | Stat. | Dyn. | Pred. |
| Offline (formal) | None | 0.531 | 0.515 | 0.741 | 0.632 | 0.297 |
| Years anonymized | None | 0.482 ( 0.049) | 0.463 ( 0.052) | 0.656 ( 0.085) | 0.576 ( 0.056) | 0.281 ( 0.016) |
| Web access 1 | Literature, data | 0.475 ( 0.056) | 0.386 ( 0.129) | 0.638 ( 0.103) | 0.572 ( 0.060) | 0.280 ( 0.017) |
| Web access 2 | Literature | 0.528 ( 0.003) | 0.406 ( 0.109) | 0.798 ( 0.057) | 0.556 ( 0.076) | 0.306 ( 0.009) |