GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash: Benchmark Head-to-Head
Three API-access models compared across FrontierMath, Terminal-Bench 2.1, DeepSWE, HLE, and Artificial Analysis indices — with per-task costs and commentator verdicts.
~4 min spoken. Keeps playing while you work in another tab.
FrontierMath: Astra's Clear Lead on Tier 4 and Tiers 1–3
GPT-6 Astra, OpenAI's API-access model, leads FrontierMath Tier 4 with 97.6% at medium effort, ahead of Claude Fable 5.1 at 87.8% at max 1. Gemini 3.8 Flash, Google DeepMind's API-access model, has no FrontierMath Tier 4 score. On FrontierMath Tiers 1–3, GPT-6 Astra again leads with 93.7% at max, followed by Claude Fable 5.1 at 90.2% at max, while Gemini 3.8 Flash has no FrontierMath Tiers 1–3 score. Bindu Reddy's read that Astra beats Fable 5.1 on math aligns with these rankings, though she qualified that Fable 5.1 remains the king of coding 8.
Data table
| Benchmark | GPT-6 Astra | Claude Fable 5.1 | Gemini 3.8 Flash |
| --- | ---: | ---: | ---: |
| FrontierMath Tier 4 | 97.6% | 87.8% | – |
| FrontierMath Tiers 1–3 | 93.7% | 90.2% | – |
| Terminal-Bench 2.1 | 89.9% | 91.4% | 87.6% |
| DeepSWE | 74.1% | – | 73.8% |
| Humanity's Last Exam | 54.7% | 59.1% | 47.8% |
| Artificial Analysis Coding Index | 77.1% | 81.6% | 76.3% |
| Artificial Analysis Intelligence Index | 54.7% | 56.8% | 47.1% |
Terminal-Bench 2.1: Fable 5.1 Leads, Astra Edges Gemini
Claude Fable 5.1 leads Terminal-Bench 2.1 at 91.4% at max, with GPT-6 Astra second at 89.9% at high and Gemini 3.8 Flash seventh at 87.6% at high 1. Chubby♨️ declared that Gemini 3.8 Flash outperforms GPT-5.6 Sol and Claude Opus 5 on Terminal-Bench 2.1, but the scores show the opposite: GPT-5.6 Sol at 89.5% and Claude Opus 5 at 89.1% both rank above Gemini 3.8 Flash at 87.6% 3.
Data table
| Rank | Model | Access | Score |
| ---: | --- | --- | ---: |
| 1 | Claude Fable 5.1 (max) | API | 91.4% |
| 2 | GPT-6 Astra (high) | API | 89.9% |
| 3 | GPT-5.6 Sol (xhigh) | API | 89.5% |
| 4 | Claude Opus 5 (max) | API | 89.1% |
| 5 | Grok 4.6 (high) | API | 88.4% |
| 6 | GPT-5.6 Terra (max) | API | 88.0% |
| 7 | Gemini 3.8 Flash (high) | API | 87.6% |
| 11 | Kimi K3 (max) | Open weights | 85.0% |
| 14 | GLM-5.3-Flash | Open weights | 84.3% |
| 25 | DeepSeek V4 Flash 0731 (max) | Open weights | 78.7% |
| 26 | DeepSeek V4 Pro 0813 (max) | Open weights | 78.7% |
| 29 | GLM-5.2 (max) | Open weights | 77.9% |
| 44 | Kimi K2.7 Code | Open weights | 67.4% |
DeepSWE: Astra Leads, Gemini Close Behind at Lower Cost
On DeepSWE, GPT-6 Astra leads with 74.1% at xhigh, Gemini 3.8 Flash follows at 73.8% at high, and Claude Fable 5.1 has no DeepSWE score 1. The cost picture matters here: GPT-6 Astra spent $6.52 per task on average, Gemini 3.8 Flash $2.36. For teams weighing software-engineering agents by what a task costs, Flash delivers nearly the same score at roughly a third of the spend.
Data table
| Model | Access | Score | $/task |
| --- | --- | ---: | ---: |
| GPT-6 Astra | API | 74.1% | $6.52 |
| Gemini 3.8 Flash | API | 73.8% | $2.36 |
| Claude Opus 5 | API | 73.6% | $12 |
| GPT-5.6 Sol | API | 72.7% | $8.39 |
| Claude Fable 5 | API | 69.9% | $13 |
| GPT-5.6 Terra | API | 69.6% | $4.95 |
| GLM-5.3 | API | 69.0% | $3.99 |
| Kimi K3 | Open weights | 68.5% | $4.65 |
| Grok 4.6 | API | 67.5% | $3.45 |
| GPT-5.6 Luna | API | 67.2% | $3.03 |
| GPT-5.5 | API | 67.0% | $7.23 |
| Gemini 3.7 Flash | API | 65.5% | $2.03 |
| GLM-5.3-Flash | Open weights | 63.4% | $0.48 |
| DeepSeek-V4-Pro | Open weights | 62.8% | $0.24 |
| Claude Opus 4.8 | API | 59.0% | $13 |
| Qwen3.8 Max | API | 57.5% | $3.73 |
| Muse Spark 1.2 | API | 54.9% | $3.70 |
| Claude Sonnet 5 | API | 53.8% | $26 |
| Grok 4.5 | API | 53.8% | $2.42 |
| DeepSeek-V4-Flash | Open weights | 53.3% | $0.10 |
| Muse Spark 1.1 | API | 53.3% | $2.36 |
| GPT-5.4 | API | 51.8% | $5.65 |
| Gemini 3.6 Flash | API | 46.7% | $4.42 |
| GLM-5.2 | Open weights | 43.8% | $3.92 |
| Gemini 3.5 Flash | API | 36.1% | $3.45 |
| Kimi K2.7 Code | Open weights | 30.5% | $2.82 |
| Claude Sonnet 4.6 | API | 29.9% | $5.52 |
| Gemini 3.1 Pro Preview | API | 11.7% | $2.14 |
Humanity's Last Exam: Fable 5.1 on Top
Claude Fable 5.1 leads Humanity's Last Exam at 59.1% at max, ahead of GPT-6 Astra at 54.7% at max, with Gemini 3.8 Flash at 47.8% at high 1. Chubby♨️ noted Gemini 3.8 Flash outperforms GPT-5.6 Sol and Claude Opus 5 on HLE, but the scores contradict that claim: GPT-5.6 Sol at 49.5% and Claude Opus 5 at 54.9% both rank above Gemini 3.8 Flash at 47.8% 3.
Artificial Analysis Coding Index: Fable 5.1 Leads
Claude Fable 5.1 tops the Artificial Analysis Coding Index at 81.6% at max, while GPT-6 Astra sits at 77.1% at high and Gemini 3.8 Flash at 76.3% at high 1. Reddy's assertion that Fable 5.1 is the king of coding matches the ranking here 8. Simon Willison's hands-on experience with Fable 5.1 producing the best SVG pelican from any Anthropic model, at a cost of $3.30, reflects the max-effort setting that drives these scores 5.
Artificial Analysis Intelligence Index: Fable 5.1 Leads Astra
The Artificial Analysis Intelligence Index, a composite measure, places Claude Fable 5.1 first at 56.8% at max, GPT-6 Astra second at 54.7% at max, and Gemini 3.8 Flash seventh at 47.1% at high 1. Reddy claimed Astra beats Fable 5.1 on reasoning, but the Intelligence Index ranking shows Fable 5.1 ahead of Astra 8.
Price and Speed Context
GPT-6 Astra costs $20.00 per 1M tokens blended with no published speed 2. Claude Fable 5.1 costs $20.00 per 1M tokens blended at 71 tokens/s. Gemini 3.8 Flash costs $1.50 per 1M tokens blended at 429 tokens/s. All three are API-access models; none are open-weight.
38 scored models without a published price are not plotted.
Data table
| Model | Access | Score | $/1M tokens |
| --- | --- | ---: | ---: |
| Claude Fable 5.1 | API | 91.4% | $20 |
| GPT-6 Astra | API | 89.9% | $20 |
| GPT-5.6 Sol | API | 89.5% | $8.00 |
| Claude Opus 5 | API | 89.1% | $10 |
| Grok 4.6 | API | 88.4% | $3.00 |
| GPT-5.6 Terra | API | 88.0% | $4.50 |
| Gemini 3.8 Flash | API | 87.6% | $1.50 |
| Qwen3.8-Flash-Next | API | 86.1% | $0.23 |
| Gemini 3.7 Flash | API | 85.8% | $1.50 |
| Muse Spark 1.3 | API | 85.8% | $2.00 |
| Kimi K3 | Open weights | 85.0% | $6.00 |
| Claude Fable 5 | API | 84.6% | $20 |
| Claude Opus 4.8 | API | 84.6% | $10 |
| GLM-5.3-Flash | Open weights | 84.3% | $0.24 |
| GPT-5.5 | API | 84.3% | $11 |
| GLM-5.3 | API | 83.9% | $2.15 |
| Claude Opus 4.7 | API | 83.1% | $10 |
| Qwen3.8 2.4T A95B | API | 82.0% | $3.00 |
| Grok 4.5 | API | 81.6% | $3.00 |
| GPT-5.6 Luna | API | 80.9% | $0.45 |
| Claude Sonnet 5 | API | 80.5% | $4.00 |
| Muse Spark 1.2 | API | 80.1% | $2.00 |
| Qwen3.8 27B | API | 79.8% | $1.13 |
| DeepSeek V4 Flash 0731 | Open weights | 78.7% | $0.66 |
| DeepSeek V4 Pro 0813 | Open weights | 78.7% | $1.98 |
| Gemini 3.5 Flash | API | 78.7% | $3.38 |
| GPT-5.4 | API | 78.3% | $5.63 |
| GLM-5.2 | Open weights | 77.9% | $2.15 |
| Muse Spark 1.1 | API | 77.9% | $2.00 |
| Gemini 3.6 Flash | API | 77.5% | $1.50 |
| Qwen3.7-Max | API | 74.5% | $3.75 |
| DeepSeek V4 Flash Vision | API | 74.2% | $0.66 |
| Claude Sonnet 4.6 | API | 71.2% | $6.00 |
| KAT Coder Pro V2 | API | 70.0% | $0.53 |
| Agnes 2.5 Pro Beta | API | 69.7% | $0.15 |
| Apodex 1.1 | API | 69.7% | $0.97 |
| Quasar 438B (max, based on GLM-5.2) | API | 69.3% | $0.90 |
| Nex-N2-Pro | API | 67.8% | $1.00 |
| Kimi K2.7 Code | Open weights | 67.4% | $1.71 |
| Agnes 2.5 Pro Alpha | API | 67.0% | $0.56 |
| Kimi K2.6 | Open weights | 65.9% | $1.71 |
| MiMo-V2.5-Pro | Open weights | 65.2% | $0.54 |
| MiniMax-M3 | Open weights | 65.2% | $0.53 |
| DeepSeek-V4-Pro | Open weights | 64.8% | $0.54 |
| Hy3 | API | 64.4% | $0.24 |
| MiMo-V2.5 | Open weights | 63.7% | $0.17 |
| DeepSeek-V4-Flash | Open weights | 61.8% | $0.12 |
| GLM-5.1 | Open weights | 61.8% | $2.00 |
| MiMo-V2-Flash | Open weights | 61.8% | $0.15 |
| Qwen3.6 Plus | API | 61.4% | $1.13 |
| Qwen3.7-Plus | API | 61.0% | $0.70 |
| GPT-5.4 Nano | API | 60.7% | $0.46 |
| Qwen3.6 27B | Open weights | 60.7% | $1.35 |
| GPT-5.4 Mini | API | 59.2% | $1.69 |
| Solar Pro 4 | API | 57.3% | $0.53 |
| Claude Sonnet 4.5 | API | 55.8% | $6.00 |
| Ling 3.0 Flash | API | 55.4% | $0.11 |
| MiniMax-M2.7 | Open weights | 55.4% | $0.53 |
| Inkling-Small | Open weights | 55.1% | $0.53 |
| Inkling | Open weights | 55.1% | $1.76 |
| Nemotron 3 Ultra 550B A55B | API | 53.9% | $1.10 |
| Gemini 3.5 Flash-Lite | API | 53.6% | $0.85 |
| GPT-5.1 | API | 52.4% | $3.44 |
| Grok Build 0.1 0616 | API | 52.1% | $1.25 |
| Muse Glimmer | API | 51.7% | $0.64 |
| Qwen3.5 397B-A17B | Open weights | 51.3% | $1.35 |
| Mistral Medium 3.5 | Open weights | 50.6% | $3.00 |
| LongCat 2.0 | API | 50.2% | $1.30 |
| GLM-4.6 | Open weights | 49.4% | $0.96 |
| Qwen3.5-122B-A10B | Open weights | 47.6% | $1.10 |
| DeepSeek-V3.2 | Open weights | 46.8% | $0.32 |
| Kimi K2.5 | Open weights | 45.7% | $1.14 |
| GLM-4.7 | Open weights | 45.3% | $1.00 |
| DeepSeek-V3.1-Terminus | Open weights | 44.9% | $0.45 |
| Qwen3.6 35B A3B | API | 44.9% | $0.84 |
| Claude Haiku 4.5 | API | 44.2% | $2.00 |
| Gemma 4 31B | API | 43.4% | $0.20 |
| Ring-2.6-1T | Open weights | 43.1% | $0.85 |
| Qwen3.5-35B-A3B | Open weights | 40.8% | $0.69 |
| Grok 4.3 | API | 39.7% | $1.56 |
| Step 3.7 Flash | Open weights | 39.3% | $0.44 |
| Gemma 4 26B A4B | Open weights | 39.0% | $0.18 |
| Nemotron 3 Super 120B A12B | API | 38.6% | $0.35 |
| qwen3-coder-next | API | 38.2% | $0.56 |
| Claude Sonnet 4 | API | 36.3% | $6.00 |
| GPT-5 | API | 35.2% | $3.44 |
| GPT-5.5 Instant (June 2026) | API | 34.8% | $11 |
| Gemini 3.1 Flash-Lite | API | 31.1% | $0.56 |
| Nova 2.0 Pro Preview | API | 29.6% | $3.44 |
| Qwen3.5-9B | Open weights | 29.2% | $0.15 |
| Gemini 2.5 Pro | API | 28.5% | $3.44 |
| Gemma 4 12B | API | 27.3% | $0.15 |
| Mercury 2 | API | 27.3% | $0.38 |
| Granite 4.2 30B | API | 26.6% | $0.28 |
| gpt-oss-120b | Open weights | 26.2% | $0.26 |
| Mistral Small 3.1 | Open weights | 26.2% | $0.15 |
| Qwen3.5-4B | Open weights | 25.8% | $0.06 |
| Nemotron 3.5 Lightning | API | 24.3% | $0.10 |
| Command A+ | Open weights | 22.8% | $4.38 |
| Mistral Small 4 | Open weights | 21.0% | $0.26 |
| Trinity Large Thinking | API | 20.6% | $0.41 |
| DeepSeek R1 (Jan '25) | API | 19.1% | $2.50 |
| Granite 4.2 8B | API | 18.4% | $0.11 |
| HyperNova 60B 2605 (high, based on gpt-oss-120b) | API | 18.4% | $0.07 |
| DeepSeek V3 (Dec '24) | API | 16.9% | $0.49 |
| Nova 2.0 Lite | API | 16.1% | $0.85 |
| DeepSeek V3 0324 | API | 13.9% | $0.48 |
| gpt-oss-20b | Open weights | 13.9% | $0.09 |
| Granite 4.2 3B | API | 13.9% | $0.05 |
| Mistral Medium 3.1 | API | 13.9% | $0.80 |
| Mistral Large 3 | Open weights | 12.0% | $0.75 |
| Qwen3 235B A22B 2507 | API | 12.0% | $0.75 |
| Solar Pro 3 | API | 12.0% | $0.26 |
| Celeris-1 | API | 11.2% | $0.33 |
| GPT-4.1 mini | API | 10.1% | $0.70 |
| Ministral 3 14B | API | 9.7% | $0.20 |
| Llama 4 Maverick | Open weights | 7.9% | $0.42 |
| Nemotron 3 Nano Omni 30B A3B Reasoning | API | 6.7% | $0.13 |
| NVIDIA Nemotron 3 Nano 30B A3B | API | 6.7% | $0.09 |
| Qwen3 Next 80B A3B | API | 6.7% | $0.41 |
| GPT-4o mini | API | 5.6% | $0.26 |
| Mistral Small 3.2 | Open weights | 5.6% | $0.15 |
| Qwen3-32B | Open weights | 5.2% | $0.28 |
| Llama 3.3 Instruct 70B | API | 4.9% | $0.67 |
| Qwen3-14B | Open weights | 4.9% | $1.31 |
| Magistral Small 1.2 | Open weights | 4.5% | $0.75 |
| o3-mini | API | 4.5% | $1.93 |
| Ministral 3 8B | API | 4.1% | $0.15 |
| GPT-4.1 nano | API | 3.7% | $0.17 |
| GPT-5 mini | API | 3.7% | $0.69 |
| Llama 4 Scout | Open weights | 3.7% | $0.31 |
| Granite 4.1 8B | API | 3.4% | $0.06 |
| Qwen3-8B | Open weights | 2.2% | $0.31 |
| Gemma 4 E4B | API | 1.9% | $0.04 |
| Llama 3.1 Instruct 8B | API | 1.5% | $0.03 |
| Qwen3 30B A3B 2507 | API | 1.5% | $0.75 |
| Ministral 3 3B | API | 0.0% | $0.10 |
Verdict by Use
For mathematical research, pick GPT-6 Astra — it leads both FrontierMath tiers, and its nearest competitor on Tier 4, Claude Fable 5.1, is nearly ten points behind at 87.8%. For agentic coding and terminal work, pick Claude Fable 5.1, provided you accept $20.00 per 1M tokens blended, since it leads Terminal-Bench 2.1, the Coding Index, and Humanity's Last Exam. For software engineering at scale on DeepSWE, pick Gemini 3.8 Flash if cost per task matters: it trails Astra on DeepSWE at $2.36 mean cost per task versus $6.52, and at $1.50 per 1M tokens blended it is the cheapest of the three. Ethan Mollick's demonstrations of Astra building a 3D Zork adaptation and procedural ocean simulations suggest it excels at complex creative coding tasks, but those are relayed anecdotes rather than benchmark scores 79.