All videos

Video briefing

LLM Evaluation and the Noise Floor: Why Public Leaderboards Cannot Choose Your Stack

GPT-6 Astra · Claude Sonnet 5 · DeepSeek-V4-Flash 3:30

Public AI rankings disguise overlapping confidence intervals and harness variance as definitive hierarchy. Here is why teams must measure their own tasks.

Transcript

What the video says

Public rankings disguise overlapping confidence intervals and harness variance as definitive hierarchy. Engineering teams must measure their own tasks to find real performance. Software teams treat public leaderboards like engineering specifications, picking models by fractional point leads. But across seventy-seven tracked benchmarks, only two publish confidence intervals, two publish cost per task, and exactly one specifies the agent harness used. Most publish point estimates that suppress experimental variance entirely.

On DeepSWE, GPT-6 Astra leads at 74.1% with an interval from 71.2 to 77.0. But nine models overlap that range, including Gemini 3.8 Flash at 73.8% and Claude Opus 5 at 73.6%. Across twenty-eight models, only three of twenty-seven adjacent rank steps are statistically separable.

When scores overlap, unit economics diverge massively. Claude Sonnet 5 reaches 53.8% at a cost of $26.40 per task. Meanwhile, DeepSeek-V4-Flash achieves 53.3% at just ten cents per task. Sonnet 5 costs 263 times more than DeepSeek-V4-Flash, despite identical statistical standing.

On Terminal-Bench 4.0, Claude Fable 5.1 leads at 57.9%, while GLM-5.3 scores 41.8% running under Anthropic's Claude Code harness. Across thirteen rows, only two of twelve adjacent steps are separable, and three models ran on foreign harnesses rather than native tooling.

Cross-suite transfer breaks down quickly. Between AIME 2025 and Terminal-Bench 2.1 across fifty-one models, correlation is 0.596, leaving 64% of variance unexplained. MiMo-V2-Flash scores 96.3% on math reasoning but 61.8% in the terminal; GPT-5 mini drops from 90.7% to 3.7%.

Across 547 benchmark pairs sharing at least thirty models, the median correlation is 0.78. More importantly, 174 pairs explain under half of each other's variance. In extreme cases like IFBench against ProofBench, correlation falls to negative 0.215, explaining only 5% of variance.

Public narratives treat leaderboard rank as an immutable ladder of raw model capability. The data contradicts this. Leaderboards measure the model and its runtime harness combined. Small point differences represent unproven noise, and high performance on public suites fails to predict real operational success.

The article's verdict is straightforward: teams must evaluate candidate models on their own workloads using their own software harnesses. Run repeated trials to establish confidence intervals. Treat overlapping bands as functional ties, and resolve those ties using real per-task costs and operational latency.

Read the complete evaluation breakdown and explore the full benchmark dataset at TerraNet Technologies.

Produced by TerraNet Technologies from the cited evidence behind the written article. Facts can change after the recorded date.