Video briefing
GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash: Frontier Math, Coding, and Cost
GPT-6 Astra · Claude Fable 5.1 · Gemini 3.8 Flash 3:59
The first week of September 2026 saw a three-way frontier-model clash between Anthropic's Claude Fable 5.1, OpenAI's GPT-6 Astra, and Google DeepMind's Gemini 3.8 Flash. We dive into the benchmarks to see where each model truly stands, comparing claims from commentators against the raw scores in frontier math, coding, and agentic engineering.
Transcript
What the video says
GPT-6 Astra, Claude Fable 5.1 and Gemini 3.8 Flash, on the benchmarks people actually cite. Independent scores, read against what the week's coverage claimed. Here is where each one leads, where it trails, and what a task costs.
September 2026 brought a three-way frontier-model clash. Anthropic shipped Claude Fable 5.1, OpenAI released GPT-6 Astra, and Google DeepMind launched Gemini 3.8 Flash. Commentators reacted immediately, with some declaring Fable 5.1 "the #1 best model in the world," while others called Flash's results a shock, claiming it outperformed top-tier models.
This overview chart shows how the three models perform across seven benchmarks. GPT-6 Astra leads in both FrontierMath tiers, while Claude Fable 5.1 takes the lead in Terminal-Bench 2.1, Humanity's Last Exam, and both Artificial Analysis indices. Gemini 3.8 Flash scores on several benchmarks, but has no scores for either FrontierMath tier or DeepSWE.
On Terminal-Bench 2.1, Claude Fable 5.1 leads with 91.4% at max effort. GPT-6 Astra follows at 89.9% at high, just ahead of GPT-5.6 Sol's 89.5%. Gemini 3.8 Flash is also present, scoring 87.6% at high, placing it seventh overall on this benchmark. This chart shows the wider competitive landscape.
This figure plots DeepSWE scores against the cost per task for various models. GPT-6 Astra leads DeepSWE at 74.1% at xhigh, but at a cost of $6.52 per task. Gemini 3.8 Flash is a close second at 73.8% at high, with a significantly lower cost of $2.36 per task. Claude Fable 5.1 has no score on DeepSWE.
This chart illustrates Terminal-Bench 2.1 scores against the price per million tokens for various models. Claude Fable 5.1 leads the benchmark at 91.4% at $20.00 per million tokens. GPT-6 Astra scores 89.9% at the same price point. Gemini 3.8 Flash, at 87.6%, offers its performance at a much lower price of $1.50 per million tokens.
Commentator claims like Fable 5.1 being "top of literally ALL AI leaderboards" are contradicted by the scores. While Fable 5.1 leads in coding and general intelligence indices, GPT-6 Astra dominates frontier math benchmarks. Similarly, claims of Gemini 3.8 Flash outperforming top models on Terminal-Bench 2.1 are not fully supported, as it trails leaders like Fable 5.1 and Astra.
For frontier mathematics, GPT-6 Astra is the clear pick. For coding and general intelligence, Claude Fable 5.1 leads. For agentic software engineering at a lower cost, Gemini 3.8 Flash is compelling, nearly matching Astra's DeepSWE score at a fraction of the cost and with superior speed. Each model has its strengths and specific use cases.
For a complete breakdown of all benchmarks, detailed scores, and deeper analysis of the models discussed, visit our website. Find out how these frontier models truly stack up and make informed decisions for your projects.
Produced by TerraNet Technologies from the cited evidence behind the written article. Facts can change after the recorded date.