Editorial illustration for LLM evaluation and the noise floor: why public leaderboards cannot choose your stack
AI analysis / Latest briefings
TerraNet Intelligence

LLM evaluation and the noise floor: why public leaderboards cannot choose your stack

Public rankings disguise overlapping confidence intervals and harness variance as definitive hierarchy. Engineering teams must measure their own tasks to find real performance.

By TerraNet Technologies7 min read10 sources
Editorial illustration for LLM evaluation and the noise floor: why public leaderboards cannot choose your stack
evaluating LLMs
LLM evaluation
LLM benchmark reliability
LLM cost per task
agent harness evaluation
Listen to this article

~7 min spoken. Keeps playing while you work in another tab.

The illusion of precision in engineering decisions

Software teams turn to public benchmark leaderboards to solve practical engineering dilemmas. A lead architect must pick a foundation model, justify the selection to leadership, and forecast operating expenditure over millions of operational turns. When a published leaderboard shows one model sitting a few percentage points above another, the natural response is to treat that hierarchy as an engineering specification: shortlist the leader, defend the decision with the published score, and build financial projections around its advertised token rates.

Yet when systems reach production, these paper advantages frequently vanish. As we documented in our briefing of 8 September 2026, user reports regularly describe capability gaps in frontier models despite their benchmark leadership 7. When downstream architectures stall or require human intervention, the procurement rationale collapses. The problem is rarely that the evaluated models changed overnight; it is that the public ranking was never statistically distinguishable in the first place.

Across the 77 benchmarks tracked in our catalogue—compiled from Datacurve and Artificial Analysis—only 2 publish run-to-run confidence intervals, 2 publish a cost per task, and exactly 1 specifies the agent harness used [[1], [4]]. The rest present precise single-point percentages that suppress experimental variance altogether. For the vast majority of public evaluations, asking whether a rank difference is mathematically real cannot even be attempted.

What the ranking does not tell you

When uncertainty is measured rigorously, published rank order disintegrates into wide overlapping bands. On DeepSWE, which reports run-to-run intervals at 95% confidence across 28 models, only 3 of 27 adjacent rank steps are statistically separable 1. The apparent sequence of superior and inferior models is largely an artifact of sorting point estimates.

DeepSWE: 3 of 27 rank steps survive their own error barsA ranked dot plot of 16 models on DeepSWE, each with its 95% run-to-run interval. 9 models' intervals overlap the leader's, and only 3 of 27 adjacent steps are statistically separable.DeepSWE: 3 of 27 rank steps survive their own error barsAmber: cannot be distinguished from the leader at 95%. Sky: separated from the leader.60%70%80%1. GPT-6 AstraGPT-6 Astra: 74.1% ±2.974.1 ±2.92. Gemini 3.8 FlashGemini 3.8 Flash: 73.8% ±1.473.8 ±1.43. Claude Opus 5Claude Opus 5: 73.6% ±3.973.6 ±3.94. GPT-5.6 SolGPT-5.6 Sol: 72.7% ±2.872.7 ±2.85. Claude Fable 5Claude Fable 5: 69.9% ±3.269.9 ±3.26. GPT-5.6 TerraGPT-5.6 Terra: 69.6% ±2.669.6 ±2.67. GLM-5.3GLM-5.3: 69.0% ±3.069.0 ±3.08. Kimi K3Kimi K3: 68.5% ±4.568.5 ±4.59. Grok 4.6Grok 4.6: 67.5% ±2.367.5 ±2.310. GPT-5.6 LunaGPT-5.6 Luna: 67.2% ±4.067.2 ±4.011. GPT-5.5GPT-5.5: 67.0% ±6.567.0 ±6.512. Gemini 3.7 FlashGemini 3.7 Flash: 65.5% ±3.165.5 ±3.113. GLM-5.3-FlashGLM-5.3-Flash: 63.4% ±4.463.4 ±4.414. DeepSeek-V4-ProDeepSeek-V4-Pro: 62.8% ±6.362.8 ±6.315. Claude Opus 4.8Claude Opus 4.8: 59.0% ±1.859.0 ±1.816. Qwen3.8 MaxQwen3.8 Max: 57.5% ±2.757.5 ±2.7leader's lower boundSource: Datacurve, DeepSWE leaderboard v1.1, https://deepswe.datacurve.ai/. Licence not stated (public leaderboard, cited with attribution).Retrieved 2026-09-05. The dot is the reported score; the bar is the 95% run-to-run interval the source publishes.Of 27 adjacent rank steps across all 28 models measured, 3 are separable at 95%. 9 models cannot be distinguished from the leader. Where twointervals overlap, the ranking between them is not evidence.
A ranked dot plot of 16 models on DeepSWE, each with its 95% run-to-run interval. 9 models' intervals overlap the leader's, and only 3 of 27 adjacent steps are statistically separable.

Showing the top 16 of 28; the counts above cover all 28.

Data table

| # | Model | Score | 95% interval |
| ---: | --- | ---: | --- |
| 1 | GPT-6 Astra | 74.1% | 71.2–77.0 |
| 2 | Gemini 3.8 Flash | 73.8% | 72.4–75.2 |
| 3 | Claude Opus 5 | 73.6% | 69.8–77.5 |
| 4 | GPT-5.6 Sol | 72.7% | 69.8–75.5 |
| 5 | Claude Fable 5 | 69.9% | 66.7–73.2 |
| 6 | GPT-5.6 Terra | 69.6% | 67.1–72.2 |
| 7 | GLM-5.3 | 69.0% | 65.9–72.0 |
| 8 | Kimi K3 | 68.5% | 64.0–73.1 |
| 9 | Grok 4.6 | 67.5% | 65.2–69.8 |
| 10 | GPT-5.6 Luna | 67.2% | 63.2–71.2 |
| 11 | GPT-5.5 | 67.0% | 60.6–73.5 |
| 12 | Gemini 3.7 Flash | 65.5% | 62.4–68.6 |
| 13 | GLM-5.3-Flash | 63.4% | 59.0–67.8 |
| 14 | DeepSeek-V4-Pro | 62.8% | 56.5–69.2 |
| 15 | Claude Opus 4.8 | 59.0% | 57.2–60.7 |
| 16 | Qwen3.8 Max | 57.5% | 54.8–60.1 |

Consider the top of the DeepSWE table. GPT-6 Astra, OpenAI's API-access model released on 3 September 2026, leads at the xhigh effort setting with 74.1% and an interval of 71.2 to 77.0 [[1], [4]]. Directly beneath it in the rankings sit Gemini 3.8 Flash from Google DeepMind at high effort with 73.8% and Claude Opus 5 from Anthropic at max effort with 73.6%. Because their intervals overlap, the ranking between them is not evidence of superior capability. In fact, 9 models overlap the leader's interval and cannot be told apart from GPT-6 Astra at 95% confidence, including GPT-5.6 Sol at max effort with 72.7% and Claude Fable 5 at xhigh effort with 69.9%.

DeepSWE against cost per task28 models plotted by DeepSWE score against cost per task.DeepSWE against cost per taskAPI accessOpen weights0%25%50%75%100%$0.10$0.30$1.00$3.00$10$30mean cost per task in USD, log scaleGPT-6 Astra · 74.1% · $6.52/1MGemini 3.8 Flash · 73.8% · $2.36/1MClaude Opus 5 · 73.6% · $12/1MGPT-5.6 Sol · 72.7% · $8.39/1MClaude Fable 5 · 69.9% · $13/1MGPT-5.6 Terra · 69.6% · $4.95/1MGLM-5.3 · 69.0% · $3.99/1MKimi K3 · 68.5% · $4.65/1MGrok 4.6 · 67.5% · $3.45/1MGPT-5.6 Luna · 67.2% · $3.03/1MGPT-5.5 · 67.0% · $7.23/1MGemini 3.7 Flash · 65.5% · $2.03/1MGLM-5.3-Flash · 63.4% · $0.48/1MDeepSeek-V4-Pro · 62.8% · $0.24/1MClaude Opus 4.8 · 59.0% · $13/1MQwen3.8 Max · 57.5% · $3.73/1MMuse Spark 1.2 · 54.9% · $3.70/1MClaude Sonnet 5 · 53.8% · $26/1MGrok 4.5 · 53.8% · $2.42/1MDeepSeek-V4-Flash · 53.3% · $0.10/1MMuse Spark 1.1 · 53.3% · $2.36/1MGPT-5.4 · 51.8% · $5.65/1MGemini 3.6 Flash · 46.7% · $4.42/1MGLM-5.2 · 43.8% · $3.92/1MGemini 3.5 Flash · 36.1% · $3.45/1MKimi K2.7 Code · 30.5% · $2.82/1MClaude Sonnet 4.6 · 29.9% · $5.52/1MGemini 3.1 Pro Preview · 11.7% · $2.14/1MGPT-6 AstraGemini 3.8 FlashClaude Opus 5GPT-5.6 TerraDeepSeek-V4-FlashDeepSeek-V4-ProSource: Datacurve, DeepSWE leaderboard v1.1, https://deepswe.datacurve.ai/. Licence not stated (public leaderboard, cited with attribution).Retrieved 2026-09-05.Cost is the mean spend per task over four runs on mini-swe-agent, as measured by Datacurve.
28 models plotted by DeepSWE score against cost per task.

Data table

| Model | Access | Score | $/task |
| --- | --- | ---: | ---: |
| GPT-6 Astra | API | 74.1% | $6.52 |
| Gemini 3.8 Flash | API | 73.8% | $2.36 |
| Claude Opus 5 | API | 73.6% | $12 |
| GPT-5.6 Sol | API | 72.7% | $8.39 |
| Claude Fable 5 | API | 69.9% | $13 |
| GPT-5.6 Terra | API | 69.6% | $4.95 |
| GLM-5.3 | API | 69.0% | $3.99 |
| Kimi K3 | Open weights | 68.5% | $4.65 |
| Grok 4.6 | API | 67.5% | $3.45 |
| GPT-5.6 Luna | API | 67.2% | $3.03 |
| GPT-5.5 | API | 67.0% | $7.23 |
| Gemini 3.7 Flash | API | 65.5% | $2.03 |
| GLM-5.3-Flash | Open weights | 63.4% | $0.48 |
| DeepSeek-V4-Pro | Open weights | 62.8% | $0.24 |
| Claude Opus 4.8 | API | 59.0% | $13 |
| Qwen3.8 Max | API | 57.5% | $3.73 |
| Muse Spark 1.2 | API | 54.9% | $3.70 |
| Claude Sonnet 5 | API | 53.8% | $26 |
| Grok 4.5 | API | 53.8% | $2.42 |
| DeepSeek-V4-Flash | Open weights | 53.3% | $0.10 |
| Muse Spark 1.1 | API | 53.3% | $2.36 |
| GPT-5.4 | API | 51.8% | $5.65 |
| Gemini 3.6 Flash | API | 46.7% | $4.42 |
| GLM-5.2 | Open weights | 43.8% | $3.92 |
| Gemini 3.5 Flash | API | 36.1% | $3.45 |
| Kimi K2.7 Code | Open weights | 30.5% | $2.82 |
| Claude Sonnet 4.6 | API | 29.9% | $5.52 |
| Gemini 3.1 Pro Preview | API | 11.7% | $2.14 |

Because these performance intervals overlap, looking at capability alone obscures massive economic divergence. On DeepSWE, Claude Sonnet 5 at max effort achieves 53.8% while averaging a task cost of $26.40 1. DeepSeek-V4-Flash, an open-weight model from DeepSeek, reaches 53.3% at max effort at a mean cost of $0.10 per task. Sonnet 5 consumes 263 times the cost of DeepSeek-V4-Flash, yet their intervals overlap completely. Similarly, DeepSeek-V4-Pro reaches 62.8% at max effort at $0.24 per task, while Sonnet 5 costs 109 times as much and shares an overlapping band. GPT-5.4 at xhigh effort posts 51.8% at $5.65 per task, running at 56 times the cost of DeepSeek-V4-Flash with overlapping bounds. Evaluating models by rank position rather than run-to-run expense leads teams to waste substantial capital on unmeasured margins.

The agent harness effect

A model score in modern autonomous benchmarks does not measure the underlying weights in isolation; it measures the model and its scaffolding together. On Terminal-Bench 4.0, which also publishes uncertainty, only 2 of 12 steps are separable across 13 rows 2. On that benchmark, a score belongs to a model and the agent harness it ran under, reported as its own distinct column.

Terminal-Bench 4.0: 2 of 12 rank steps survive their own error barsA ranked dot plot of 13 models on Terminal-Bench 4.0, each with its 95% run-to-run interval. 2 models' intervals overlap the leader's, and only 2 of 12 adjacent steps are statistically separable.Terminal-Bench 4.0: 2 of 12 rank steps survive their own error barsAmber: cannot be distinguished from the leader at 95%. Sky: separated from the leader.10%20%30%40%50%60%1. Claude Fable 5.1Claude Fable 5.1: 57.9% ±3.857.9 ±3.82. Claude Opus 5Claude Opus 5: 51.8% ±3.451.8 ±3.43. Claude Fable 5Claude Fable 5: 44.5% ±3.944.5 ±3.94. GLM-5.3GLM-5.3: 41.8% ±3.241.8 ±3.25. GPT-5.6 SolGPT-5.6 Sol: 37.3% ±3.837.3 ±3.86. Claude Opus 4.8Claude Opus 4.8: 23.6% ±3.623.6 ±3.67. GPT-5.6 TerraGPT-5.6 Terra: 21.5% ±3.321.5 ±3.38. Grok 4.6Grok 4.6: 20.3% ±3.120.3 ±3.19. Gemini 3.8 FlashGemini 3.8 Flash: 19.1% ±3.419.1 ±3.410. GPT-5.6 LunaGPT-5.6 Luna: 17.3% ±2.917.3 ±2.911. Claude Sonnet 5Claude Sonnet 5: 12.4% ±3.112.4 ±3.112. Grok 4.5Grok 4.5: 12.4% ±2.612.4 ±2.613. Gemini 3.7 FlashGemini 3.7 Flash: 11.2% ±2.511.2 ±2.5leader's lower boundSource: Terminal-Bench 4.0, Stanford / Harbor / Laude Institute, https://www.tbench.ai/leaderboard. Licence Apache-2.0. Retrieved 2026-09-10. Thedot is the reported score; the bar is the 95% run-to-run interval the source publishes.Of 12 adjacent rank steps across all 13 models measured, 2 are separable at 95%. 2 models cannot be distinguished from the leader. Where twointervals overlap, the ranking between them is not evidence.
A ranked dot plot of 13 models on Terminal-Bench 4.0, each with its 95% run-to-run interval. 2 models' intervals overlap the leader's, and only 2 of 12 adjacent steps are statistically separable.

Data table

| # | Model | Score | 95% interval |
| ---: | --- | ---: | --- |
| 1 | Claude Fable 5.1 | 57.9% | 54.1–61.6 |
| 2 | Claude Opus 5 | 51.8% | 48.4–55.2 |
| 3 | Claude Fable 5 | 44.5% | 40.7–48.4 |
| 4 | GLM-5.3 | 41.8% | 38.6–45.1 |
| 5 | GPT-5.6 Sol | 37.3% | 33.5–41.1 |
| 6 | Claude Opus 4.8 | 23.6% | 20.1–27.2 |
| 7 | GPT-5.6 Terra | 21.5% | 18.3–24.8 |
| 8 | Grok 4.6 | 20.3% | 17.2–23.4 |
| 9 | Gemini 3.8 Flash | 19.1% | 15.7–22.4 |
| 10 | GPT-5.6 Luna | 17.3% | 14.4–20.1 |
| 11 | Claude Sonnet 5 | 12.4% | 9.4–15.5 |
| 12 | Grok 4.5 | 12.4% | 9.8–15.0 |
| 13 | Gemini 3.7 Flash | 11.2% | 8.8–13.7 |

Crucially, 3 of 13 rows on Terminal-Bench 4.0 used an agent harness from an organisation different from the model creator 2. Gemini 3.7 Flash at high effort and Gemini 3.8 Flash at high effort were evaluated under the mini-SWE-agent harness, scoring 11.2% and 19.1% respectively. Meanwhile, GLM-5.3, an API-access model from Zhipu AI, achieved 41.8% at max effort while running under Anthropic's Claude Code harness. A public leaderboard that mixes agent runtimes is not comparing models like with like, and attributing the resulting performance differences to the model weights misidentifies the source of system behaviour.

Benchmark transfer and predictive limits

A team might hope that an aggregate leaderboard position indicates broad, transferable capability. The cross-benchmark data proves otherwise. Across 547 pairs of benchmarks sharing at least 30 models, the median correlation is 0.78, and 174 pairs explain under half of each other's variance 4.

AIME 2025 predicts 36% of Terminal-Bench 2.1A scatter plot of 51 models by AIME 2025 score horizontally and Terminal-Bench 2.1 score vertically, correlation r = 0.60.AIME 2025 predicts 36% of Terminal-Bench 2.10%0%25%25%50%50%75%75%100%100%AIME 2025 →Terminal-Bench 2.1 →Claude Haiku 4.5: 83.7% on AIME 2025, 44.2% on Terminal-Bench 2.1Claude Sonnet 4.5: 88.0% on AIME 2025, 55.8% on Terminal-Bench 2.1Claude Sonnet 4: 74.3% on AIME 2025, 36.3% on Terminal-Bench 2.1Command A: 13.0% on AIME 2025, 22.8% on Terminal-Bench 2.1DeepSeek R1 (Jan '25): 68.0% on AIME 2025, 19.1% on Terminal-Bench 2.1DeepSeek V3 0324: 41.0% on AIME 2025, 13.9% on Terminal-Bench 2.1DeepSeek-V3.1-Terminus: 89.7% on AIME 2025, 44.9% on Terminal-Bench 2.1DeepSeek-V3.2: 92.0% on AIME 2025, 46.8% on Terminal-Bench 2.1DeepSeek V3 (Dec '24): 26.0% on AIME 2025, 16.9% on Terminal-Bench 2.1Devstral 2: 36.7% on AIME 2025, 30.3% on Terminal-Bench 2.1Devstral Small 2: 34.3% on AIME 2025, 29.6% on Terminal-Bench 2.1Gemini 2.5 Pro: 87.7% on AIME 2025, 28.5% on Terminal-Bench 2.1Gemma 3 12B Instruct: 18.3% on AIME 2025, 0.0% on Terminal-Bench 2.1Gemma 3 27B Instruct: 20.7% on AIME 2025, 4.5% on Terminal-Bench 2.1Gemma 3 4B Instruct: 12.7% on AIME 2025, 0.4% on Terminal-Bench 2.1Gemma 3n E4B Instruct: 14.3% on AIME 2025, 0.7% on Terminal-Bench 2.1GLM-4.6: 86.0% on AIME 2025, 49.4% on Terminal-Bench 2.1GLM-4.7: 95.0% on AIME 2025, 45.3% on Terminal-Bench 2.1GPT-4.1 mini: 46.3% on AIME 2025, 10.1% on Terminal-Bench 2.1GPT-4.1 nano: 24.0% on AIME 2025, 3.7% on Terminal-Bench 2.1GPT-4o mini: 14.7% on AIME 2025, 5.6% on Terminal-Bench 2.1GPT-5.1: 94.0% on AIME 2025, 52.4% on Terminal-Bench 2.1GPT-5 mini: 90.7% on AIME 2025, 3.7% on Terminal-Bench 2.1GPT-5: 94.3% on AIME 2025, 35.2% on Terminal-Bench 2.1gpt-oss-120b: 93.4% on AIME 2025, 26.2% on Terminal-Bench 2.1gpt-oss-20b: 89.3% on AIME 2025, 13.9% on Terminal-Bench 2.1K-EXAONE: 90.3% on AIME 2025, 30.3% on Terminal-Bench 2.1Llama 3.1 Instruct 8B: 4.3% on AIME 2025, 1.5% on Terminal-Bench 2.1Llama 3.3 Instruct 70B: 7.7% on AIME 2025, 4.9% on Terminal-Bench 2.1Llama 4 Maverick: 19.3% on AIME 2025, 7.9% on Terminal-Bench 2.1Llama 4 Scout: 14.0% on AIME 2025, 3.7% on Terminal-Bench 2.1Magistral Medium 1.2: 82.0% on AIME 2025, 12.4% on Terminal-Bench 2.1Magistral Small 1.2: 80.3% on AIME 2025, 4.5% on Terminal-Bench 2.1MiMo-V2-Flash: 96.3% on AIME 2025, 61.8% on Terminal-Bench 2.1Ministral 3 14B: 30.0% on AIME 2025, 9.7% on Terminal-Bench 2.1Ministral 3 3B: 22.0% on AIME 2025, 0.0% on Terminal-Bench 2.1Ministral 3 8B: 31.7% on AIME 2025, 4.1% on Terminal-Bench 2.1Mistral Large 3: 38.0% on AIME 2025, 12.0% on Terminal-Bench 2.1Mistral Medium 3.1: 38.3% on AIME 2025, 13.9% on Terminal-Bench 2.1Mistral Small 3.1: 3.7% on AIME 2025, 26.2% on Terminal-Bench 2.1Mistral Small 3.2: 27.0% on AIME 2025, 5.6% on Terminal-Bench 2.1Nova 2.0 Lite: 94.3% on AIME 2025, 16.1% on Terminal-Bench 2.1Nova 2.0 Pro Preview: 89.0% on AIME 2025, 29.6% on Terminal-Bench 2.1NVIDIA Nemotron 3 Nano 30B A3B: 91.0% on AIME 2025, 6.7% on Terminal-Bench 2.1Phi-4 Mini Instruct: 6.7% on AIME 2025, 0.4% on Terminal-Bench 2.1Qwen3-14B: 58.0% on AIME 2025, 4.9% on Terminal-Bench 2.1Qwen3 235B A22B 2507: 91.0% on AIME 2025, 12.0% on Terminal-Bench 2.1Qwen3 30B A3B 2507: 56.3% on AIME 2025, 1.5% on Terminal-Bench 2.1Qwen3-32B: 73.0% on AIME 2025, 5.2% on Terminal-Bench 2.1Qwen3-8B: 24.3% on AIME 2025, 2.2% on Terminal-Bench 2.1Qwen3 Next 80B A3B: 84.3% on AIME 2025, 6.7% on Terminal-Bench 2.1MiMo-V2-FlashClaude Sonnet 4.5GPT-5 miniNVIDIA Nemotron 3 Nano 30…Source: Artificial Analysis, https://artificialanalysis.ai/ Licence Free API, attribution required. Retrieved 2026-09-05. Each point is onemodel's best published score on both.r = 0.596, so AIME 2025 explains 36% of the variance in Terminal-Bench 2.1 across these 51 models, and 64% of it is something else. The twobenchmarks share no inputs.
A scatter plot of 51 models by AIME 2025 score horizontally and Terminal-Bench 2.1 score vertically, correlation r = 0.60.

Data table

| Model | AIME 2025 | Terminal-Bench 2.1 |
| --- | ---: | ---: |
| MiMo-V2-Flash | 96.3% | 61.8% |
| GLM-4.7 | 95.0% | 45.3% |
| GPT-5 | 94.3% | 35.2% |
| Nova 2.0 Lite | 94.3% | 16.1% |
| GPT-5.1 | 94.0% | 52.4% |
| gpt-oss-120b | 93.4% | 26.2% |
| DeepSeek-V3.2 | 92.0% | 46.8% |
| NVIDIA Nemotron 3 Nano 30B A3B | 91.0% | 6.7% |
| Qwen3 235B A22B 2507 | 91.0% | 12.0% |
| GPT-5 mini | 90.7% | 3.7% |
| K-EXAONE | 90.3% | 30.3% |
| DeepSeek-V3.1-Terminus | 89.7% | 44.9% |
| gpt-oss-20b | 89.3% | 13.9% |
| Nova 2.0 Pro Preview | 89.0% | 29.6% |
| Claude Sonnet 4.5 | 88.0% | 55.8% |
| Gemini 2.5 Pro | 87.7% | 28.5% |
| GLM-4.6 | 86.0% | 49.4% |
| Qwen3 Next 80B A3B | 84.3% | 6.7% |
| Claude Haiku 4.5 | 83.7% | 44.2% |
| Magistral Medium 1.2 | 82.0% | 12.4% |

A comparison between AIME 2025 and Terminal-Bench 2.1 across 51 models demonstrates this disconnect 4. The correlation between the two suites is r = 0.596. While this relationship explains 36% of the variance, it leaves 64% completely unexplained. High competence on competition mathematics offers little guarantee of competency in an interactive terminal. MiMo-V2-Flash reaches 96.3% at reasoning on AIME 2025 but drops to 61.8% at non-reasoning on Terminal-Bench 2.1; GPT-5 mini reaches 90.7% at high effort on AIME 2025 but manages only 3.7% at high effort in the terminal environment.

174 of 547 benchmark pairs explain under half of each otherA histogram of 547 benchmark pairs by how much of one benchmark's variance the other explains, from 0 to 100 per cent, with 174 pairs below half.174 of 547 benchmark pairs explain under half of each otherEach pair counted once. Higher is not better - it means the two measure the same thing.0204060801006 pair(s) between 0% and 10%6015 pair(s) between 10% and 20%151033 pair(s) between 20% and 30%332042 pair(s) between 30% and 40%423036%78 pair(s) between 40% and 50%784092 pair(s) between 50% and 60%9250100 pair(s) between 60% and 70%1006090 pair(s) between 70% and 80%907083 pair(s) between 80% and 90%83808 pair(s) between 90% and 100%890Percentage of one benchmark’s variance explained by the other →Source: Epoch AI and Artificial Analysis, joined by this site. Licence CC BY 4.0 and attribution respectively. Retrieved 2026-09-04. Every pair ofbenchmarks with at least 30 models scored on both, 547 in total; the bar is how many pairs fall in that band.174 of 547 pairs explain under half the variance in each other. A pair near 100% is more often a composite index against one of its own componentsthan two benchmarks that genuinely agree.
A histogram of 547 benchmark pairs by how much of one benchmark's variance the other explains, from 0 to 100 per cent, with 174 pairs below half.

Data table

| Variance explained | Pairs |
| --- | ---: |
| 0–10% | 6 |
| 10–20% | 15 |
| 20–30% | 33 |
| 30–40% | 42 |
| 40–50% | 78 |
| 50–60% | 92 |
| 60–70% | 100 |
| 70–80% | 90 |
| 80–90% | 83 |
| 90–100% | 8 |

In extreme pairings, transfer disappears entirely. Artificial Analysis's IFBench against ProofBench across 33 models yields an r of -0.215, explaining only 5% of variance 4. Comparing τ²-Bench against ProofBench across 33 models produces an r of 0.172, explaining just 3%. IFBench against Mystery Game Puzzles across 36 models shows an r of 0.193, explaining 4%, and against Terminal Bench across 34 models shows an r of 0.251, explaining 6%. Relying on a third-party composite to predict in-house success is mathematically indefensible.

What a team can do instead

Because public leaderboards evaluate external tasks under external scaffolding, an organisation must evaluate models on its own operational workloads using its own software harness. A public benchmark tests someone else's problem under someone else's agent environment; our cross-benchmark transfer data demonstrates how weakly those conditions predict alternate operational goals.

However, building an internal evaluation suite exposes teams to the same statistical traps observed in public data. If a benchmark running hundreds of trials per model can separate only a handful of its own rank steps, a dozen hand-written internal tasks run once will separate nothing. A team that creates a small test suite, runs each candidate model once, and ranks them by absolute score has simply reproduced the public leaderboard error inside their own infrastructure.

The case for an in-house evaluation suite is not that it is inherently more precise; it is that it measures the right thing. Precision comes strictly from repeated trials and explicit confidence bounds. Relevance comes from choosing the tasks and the harness that match production architecture. A team should construct a tightly scoped set of representative enterprise tasks, execute dozens of runs per model configuration, compute the resulting variance, and treat overlapping ranges as ties to be decided by unit economics and execution latency.

What would settle it

Resolving these ambiguities requires reporting that treats evaluation as an empirical measurement rather than a tournament. The missing metric across the ecosystem is systematic run-to-run variance, paired with explicit declarations of runtime harness dependencies and real per-task execution costs. Primary benchmark creators must publish standard errors across multi-trial evaluations as a baseline requirement.

Where third-party aggregators relay benchmark results, they must stop dropping the underlying uncertainty. When a benchmark source measures variance but the reporting aggregator publishes only a single stripped point score, the supply chain strips engineers of the statistical context needed to make informed decisions. Until platforms preserve these intervals and publish run-to-run variance across native harness environments, engineering teams must treat every public rank step as unproven noise.

AI Tools

    LLM evaluation and the noise floor: why public leaderboards cannot choose your stack | TerraNet Technologies