LLM evaluation and the noise floor: why public leaderboards cannot choose your stack
Public rankings disguise overlapping confidence intervals and harness variance as definitive hierarchy. Engineering teams must measure their own tasks to find real performance.
~7 min spoken. Keeps playing while you work in another tab.
The illusion of precision in engineering decisions
Software teams turn to public benchmark leaderboards to solve practical engineering dilemmas. A lead architect must pick a foundation model, justify the selection to leadership, and forecast operating expenditure over millions of operational turns. When a published leaderboard shows one model sitting a few percentage points above another, the natural response is to treat that hierarchy as an engineering specification: shortlist the leader, defend the decision with the published score, and build financial projections around its advertised token rates.
Yet when systems reach production, these paper advantages frequently vanish. As we documented in our briefing of 8 September 2026, user reports regularly describe capability gaps in frontier models despite their benchmark leadership 7. When downstream architectures stall or require human intervention, the procurement rationale collapses. The problem is rarely that the evaluated models changed overnight; it is that the public ranking was never statistically distinguishable in the first place.
Across the 77 benchmarks tracked in our catalogue—compiled from Datacurve and Artificial Analysis—only 2 publish run-to-run confidence intervals, 2 publish a cost per task, and exactly 1 specifies the agent harness used [[1], [4]]. The rest present precise single-point percentages that suppress experimental variance altogether. For the vast majority of public evaluations, asking whether a rank difference is mathematically real cannot even be attempted.
What the ranking does not tell you
When uncertainty is measured rigorously, published rank order disintegrates into wide overlapping bands. On DeepSWE, which reports run-to-run intervals at 95% confidence across 28 models, only 3 of 27 adjacent rank steps are statistically separable 1. The apparent sequence of superior and inferior models is largely an artifact of sorting point estimates.
Showing the top 16 of 28; the counts above cover all 28.
Data table
| # | Model | Score | 95% interval |
| ---: | --- | ---: | --- |
| 1 | GPT-6 Astra | 74.1% | 71.2–77.0 |
| 2 | Gemini 3.8 Flash | 73.8% | 72.4–75.2 |
| 3 | Claude Opus 5 | 73.6% | 69.8–77.5 |
| 4 | GPT-5.6 Sol | 72.7% | 69.8–75.5 |
| 5 | Claude Fable 5 | 69.9% | 66.7–73.2 |
| 6 | GPT-5.6 Terra | 69.6% | 67.1–72.2 |
| 7 | GLM-5.3 | 69.0% | 65.9–72.0 |
| 8 | Kimi K3 | 68.5% | 64.0–73.1 |
| 9 | Grok 4.6 | 67.5% | 65.2–69.8 |
| 10 | GPT-5.6 Luna | 67.2% | 63.2–71.2 |
| 11 | GPT-5.5 | 67.0% | 60.6–73.5 |
| 12 | Gemini 3.7 Flash | 65.5% | 62.4–68.6 |
| 13 | GLM-5.3-Flash | 63.4% | 59.0–67.8 |
| 14 | DeepSeek-V4-Pro | 62.8% | 56.5–69.2 |
| 15 | Claude Opus 4.8 | 59.0% | 57.2–60.7 |
| 16 | Qwen3.8 Max | 57.5% | 54.8–60.1 |
Consider the top of the DeepSWE table. GPT-6 Astra, OpenAI's API-access model released on 3 September 2026, leads at the xhigh effort setting with 74.1% and an interval of 71.2 to 77.0 [[1], [4]]. Directly beneath it in the rankings sit Gemini 3.8 Flash from Google DeepMind at high effort with 73.8% and Claude Opus 5 from Anthropic at max effort with 73.6%. Because their intervals overlap, the ranking between them is not evidence of superior capability. In fact, 9 models overlap the leader's interval and cannot be told apart from GPT-6 Astra at 95% confidence, including GPT-5.6 Sol at max effort with 72.7% and Claude Fable 5 at xhigh effort with 69.9%.
Data table
| Model | Access | Score | $/task |
| --- | --- | ---: | ---: |
| GPT-6 Astra | API | 74.1% | $6.52 |
| Gemini 3.8 Flash | API | 73.8% | $2.36 |
| Claude Opus 5 | API | 73.6% | $12 |
| GPT-5.6 Sol | API | 72.7% | $8.39 |
| Claude Fable 5 | API | 69.9% | $13 |
| GPT-5.6 Terra | API | 69.6% | $4.95 |
| GLM-5.3 | API | 69.0% | $3.99 |
| Kimi K3 | Open weights | 68.5% | $4.65 |
| Grok 4.6 | API | 67.5% | $3.45 |
| GPT-5.6 Luna | API | 67.2% | $3.03 |
| GPT-5.5 | API | 67.0% | $7.23 |
| Gemini 3.7 Flash | API | 65.5% | $2.03 |
| GLM-5.3-Flash | Open weights | 63.4% | $0.48 |
| DeepSeek-V4-Pro | Open weights | 62.8% | $0.24 |
| Claude Opus 4.8 | API | 59.0% | $13 |
| Qwen3.8 Max | API | 57.5% | $3.73 |
| Muse Spark 1.2 | API | 54.9% | $3.70 |
| Claude Sonnet 5 | API | 53.8% | $26 |
| Grok 4.5 | API | 53.8% | $2.42 |
| DeepSeek-V4-Flash | Open weights | 53.3% | $0.10 |
| Muse Spark 1.1 | API | 53.3% | $2.36 |
| GPT-5.4 | API | 51.8% | $5.65 |
| Gemini 3.6 Flash | API | 46.7% | $4.42 |
| GLM-5.2 | Open weights | 43.8% | $3.92 |
| Gemini 3.5 Flash | API | 36.1% | $3.45 |
| Kimi K2.7 Code | Open weights | 30.5% | $2.82 |
| Claude Sonnet 4.6 | API | 29.9% | $5.52 |
| Gemini 3.1 Pro Preview | API | 11.7% | $2.14 |
Because these performance intervals overlap, looking at capability alone obscures massive economic divergence. On DeepSWE, Claude Sonnet 5 at max effort achieves 53.8% while averaging a task cost of $26.40 1. DeepSeek-V4-Flash, an open-weight model from DeepSeek, reaches 53.3% at max effort at a mean cost of $0.10 per task. Sonnet 5 consumes 263 times the cost of DeepSeek-V4-Flash, yet their intervals overlap completely. Similarly, DeepSeek-V4-Pro reaches 62.8% at max effort at $0.24 per task, while Sonnet 5 costs 109 times as much and shares an overlapping band. GPT-5.4 at xhigh effort posts 51.8% at $5.65 per task, running at 56 times the cost of DeepSeek-V4-Flash with overlapping bounds. Evaluating models by rank position rather than run-to-run expense leads teams to waste substantial capital on unmeasured margins.
The agent harness effect
A model score in modern autonomous benchmarks does not measure the underlying weights in isolation; it measures the model and its scaffolding together. On Terminal-Bench 4.0, which also publishes uncertainty, only 2 of 12 steps are separable across 13 rows 2. On that benchmark, a score belongs to a model and the agent harness it ran under, reported as its own distinct column.
Data table
| # | Model | Score | 95% interval |
| ---: | --- | ---: | --- |
| 1 | Claude Fable 5.1 | 57.9% | 54.1–61.6 |
| 2 | Claude Opus 5 | 51.8% | 48.4–55.2 |
| 3 | Claude Fable 5 | 44.5% | 40.7–48.4 |
| 4 | GLM-5.3 | 41.8% | 38.6–45.1 |
| 5 | GPT-5.6 Sol | 37.3% | 33.5–41.1 |
| 6 | Claude Opus 4.8 | 23.6% | 20.1–27.2 |
| 7 | GPT-5.6 Terra | 21.5% | 18.3–24.8 |
| 8 | Grok 4.6 | 20.3% | 17.2–23.4 |
| 9 | Gemini 3.8 Flash | 19.1% | 15.7–22.4 |
| 10 | GPT-5.6 Luna | 17.3% | 14.4–20.1 |
| 11 | Claude Sonnet 5 | 12.4% | 9.4–15.5 |
| 12 | Grok 4.5 | 12.4% | 9.8–15.0 |
| 13 | Gemini 3.7 Flash | 11.2% | 8.8–13.7 |
Crucially, 3 of 13 rows on Terminal-Bench 4.0 used an agent harness from an organisation different from the model creator 2. Gemini 3.7 Flash at high effort and Gemini 3.8 Flash at high effort were evaluated under the mini-SWE-agent harness, scoring 11.2% and 19.1% respectively. Meanwhile, GLM-5.3, an API-access model from Zhipu AI, achieved 41.8% at max effort while running under Anthropic's Claude Code harness. A public leaderboard that mixes agent runtimes is not comparing models like with like, and attributing the resulting performance differences to the model weights misidentifies the source of system behaviour.
Benchmark transfer and predictive limits
A team might hope that an aggregate leaderboard position indicates broad, transferable capability. The cross-benchmark data proves otherwise. Across 547 pairs of benchmarks sharing at least 30 models, the median correlation is 0.78, and 174 pairs explain under half of each other's variance 4.
Data table
| Model | AIME 2025 | Terminal-Bench 2.1 |
| --- | ---: | ---: |
| MiMo-V2-Flash | 96.3% | 61.8% |
| GLM-4.7 | 95.0% | 45.3% |
| GPT-5 | 94.3% | 35.2% |
| Nova 2.0 Lite | 94.3% | 16.1% |
| GPT-5.1 | 94.0% | 52.4% |
| gpt-oss-120b | 93.4% | 26.2% |
| DeepSeek-V3.2 | 92.0% | 46.8% |
| NVIDIA Nemotron 3 Nano 30B A3B | 91.0% | 6.7% |
| Qwen3 235B A22B 2507 | 91.0% | 12.0% |
| GPT-5 mini | 90.7% | 3.7% |
| K-EXAONE | 90.3% | 30.3% |
| DeepSeek-V3.1-Terminus | 89.7% | 44.9% |
| gpt-oss-20b | 89.3% | 13.9% |
| Nova 2.0 Pro Preview | 89.0% | 29.6% |
| Claude Sonnet 4.5 | 88.0% | 55.8% |
| Gemini 2.5 Pro | 87.7% | 28.5% |
| GLM-4.6 | 86.0% | 49.4% |
| Qwen3 Next 80B A3B | 84.3% | 6.7% |
| Claude Haiku 4.5 | 83.7% | 44.2% |
| Magistral Medium 1.2 | 82.0% | 12.4% |
A comparison between AIME 2025 and Terminal-Bench 2.1 across 51 models demonstrates this disconnect 4. The correlation between the two suites is r = 0.596. While this relationship explains 36% of the variance, it leaves 64% completely unexplained. High competence on competition mathematics offers little guarantee of competency in an interactive terminal. MiMo-V2-Flash reaches 96.3% at reasoning on AIME 2025 but drops to 61.8% at non-reasoning on Terminal-Bench 2.1; GPT-5 mini reaches 90.7% at high effort on AIME 2025 but manages only 3.7% at high effort in the terminal environment.
Data table
| Variance explained | Pairs |
| --- | ---: |
| 0–10% | 6 |
| 10–20% | 15 |
| 20–30% | 33 |
| 30–40% | 42 |
| 40–50% | 78 |
| 50–60% | 92 |
| 60–70% | 100 |
| 70–80% | 90 |
| 80–90% | 83 |
| 90–100% | 8 |
In extreme pairings, transfer disappears entirely. Artificial Analysis's IFBench against ProofBench across 33 models yields an r of -0.215, explaining only 5% of variance 4. Comparing τ²-Bench against ProofBench across 33 models produces an r of 0.172, explaining just 3%. IFBench against Mystery Game Puzzles across 36 models shows an r of 0.193, explaining 4%, and against Terminal Bench across 34 models shows an r of 0.251, explaining 6%. Relying on a third-party composite to predict in-house success is mathematically indefensible.
What a team can do instead
Because public leaderboards evaluate external tasks under external scaffolding, an organisation must evaluate models on its own operational workloads using its own software harness. A public benchmark tests someone else's problem under someone else's agent environment; our cross-benchmark transfer data demonstrates how weakly those conditions predict alternate operational goals.
However, building an internal evaluation suite exposes teams to the same statistical traps observed in public data. If a benchmark running hundreds of trials per model can separate only a handful of its own rank steps, a dozen hand-written internal tasks run once will separate nothing. A team that creates a small test suite, runs each candidate model once, and ranks them by absolute score has simply reproduced the public leaderboard error inside their own infrastructure.
The case for an in-house evaluation suite is not that it is inherently more precise; it is that it measures the right thing. Precision comes strictly from repeated trials and explicit confidence bounds. Relevance comes from choosing the tasks and the harness that match production architecture. A team should construct a tightly scoped set of representative enterprise tasks, execute dozens of runs per model configuration, compute the resulting variance, and treat overlapping ranges as ties to be decided by unit economics and execution latency.
What would settle it
Resolving these ambiguities requires reporting that treats evaluation as an empirical measurement rather than a tournament. The missing metric across the ecosystem is systematic run-to-run variance, paired with explicit declarations of runtime harness dependencies and real per-task execution costs. Primary benchmark creators must publish standard errors across multi-trial evaluations as a baseline requirement.
Where third-party aggregators relay benchmark results, they must stop dropping the underlying uncertainty. When a benchmark source measures variance but the reporting aggregator publishes only a single stripped point score, the supply chain strips engineers of the statistical context needed to make informed decisions. Until platforms preserve these intervals and publish run-to-run variance across native harness environments, engineering teams must treat every public rank step as unproven noise.