← SciExam for ENSO

Detailed results

12 LLM agents, each given 360 minutes to build a stochastic conceptual ENSO model · one frozen grader · statistics against the observed Niño3 / Niño4 record, 1950-2019

Leaderboard model equations

Rank by

#ModelCompositeStat Dyn. consistencyPredComposite by checkpoint

Composite = 0.3 × statistics + 0.3 × dynamical consistency + 0.4 × predictivity, behind a validity gate (a 50 yr free run must stay finite). Statistics: how well the model's long free run reproduces the observed PDF, seasonal variance, autocorrelation, event frequencies and El Niño diversity. Dynamical consistency: how much of the unobserved variables an ensemble Kalman smoother recovers through the model (mean NSE). Predictivity: forecast skill against persistence on the held-out years 2015-2024. The reference model is the human-built model of 2402.04585, graded by the same grader; it sits in the table at its own score as a divider and is not ranked. Statistics uses the revised rule of 2026-09-16 (formulas under the Statistics tab); physics and predictivity are unchanged. Each submission is evaluated under the time convention stated by its agent, because the interface did not fix which calendar month corresponds to model time.

Supplementary runs (below the line, not ranked): the same agent as Claude Fable 5.1 with one condition changed, graded by the same frozen grader. The two web runs had a different task text that required consulting the literature through a web search tool (one with literature only, one with literature and data); the literature-only run is valid with an exception: a search listing exposed the prohibited reference's metadata and abstract, and no use of it was observed. The years anonymised run replaced calendar years with relative labels (months kept). Each is a single run.

observation (1980-2019, 40 yr)