AI analysis / Latest briefings
TerraNet Intelligence

GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash: Benchmark Head-to-Head

Three API-access models compared across FrontierMath, Terminal-Bench 2.1, DeepSWE, HLE, and Artificial Analysis indices — with per-task costs and commentator verdicts.

By TerraNet Technologies4 min read10 sources
GPT-6 Astra benchmarks
Claude Fable 5.1 vs Gemini 3.8 Flash
FrontierMath Tier 4 comparison
DeepSWE cost per task
Terminal-Bench 2.1 ranking
Artificial Analysis Coding Index
Listen to this article

~4 min spoken. Keeps playing while you work in another tab.

FrontierMath: Astra's Clear Lead on Tier 4 and Tiers 1–3

GPT-6 Astra, OpenAI's API-access model, leads FrontierMath Tier 4 with 97.6% at medium effort, ahead of Claude Fable 5.1 at 87.8% at max 1. Gemini 3.8 Flash, Google DeepMind's API-access model, has no FrontierMath Tier 4 score. On FrontierMath Tiers 1–3, GPT-6 Astra again leads with 93.7% at max, followed by Claude Fable 5.1 at 90.2% at max, while Gemini 3.8 Flash has no FrontierMath Tiers 1–3 score. Bindu Reddy's read that Astra beats Fable 5.1 on math aligns with these rankings, though she qualified that Fable 5.1 remains the king of coding 8.

GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash across 7 benchmarksGPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash across 7 benchmarks.GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash across 7 benchmarksGPT-6 AstraClaude Fable 5.1Gemini 3.8 Flash0%25%50%75%100%GPT-6 Astra · FrontierMath Tier 4 · 97.6%98Claude Fable 5.1 · FrontierMath Tier 4 · 87.8%88FrontierMath Tier 4GPT-6 Astra · FrontierMath Tiers 1–3 · 93.7%94Claude Fable 5.1 · FrontierMath Tiers 1–3 · 90.2%90FrontierMath Tiers1–3GPT-6 Astra · Terminal-Bench 2.1 · 89.9%90Claude Fable 5.1 · Terminal-Bench 2.1 · 91.4%91Gemini 3.8 Flash · Terminal-Bench 2.1 · 87.6%88Terminal-Bench 2.1GPT-6 Astra · DeepSWE · 74.1%74Gemini 3.8 Flash · DeepSWE · 73.8%74DeepSWEGPT-6 Astra · Humanity's Last Exam · 54.7%55Claude Fable 5.1 · Humanity's Last Exam · 59.1%59Gemini 3.8 Flash · Humanity's Last Exam · 47.8%48Humanity's Last ExamGPT-6 Astra · Artificial Analysis Coding Index · 77.1%77Claude Fable 5.1 · Artificial Analysis Coding Index · 81.6%82Gemini 3.8 Flash · Artificial Analysis Coding Index · 76.3%76Artificial AnalysisCoding IndexGPT-6 Astra · Artificial Analysis Intelligence Index · 54.7%55Claude Fable 5.1 · Artificial Analysis Intelligence Index · 56.8%57Gemini 3.8 Flash · Artificial Analysis Intelligence Index · 47.1%47Artificial AnalysisIntelligence IndexSource: Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks' [online resource]. Licence CCBY 4.0. Retrieved 2026-09-04.Source: Artificial Analysis, https://artificialanalysis.ai/ Licence Free API, attribution required. Retrieved 2026-09-04.Source: Datacurve, DeepSWE leaderboard v1.1, https://deepswe.datacurve.ai/. Licence not stated (public leaderboard, cited with attribution).Retrieved 2026-09-04.
GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash across 7 benchmarks.

Data table

| Benchmark | GPT-6 Astra | Claude Fable 5.1 | Gemini 3.8 Flash |
| --- | ---: | ---: | ---: |
| FrontierMath Tier 4 | 97.6% | 87.8% | – |
| FrontierMath Tiers 1–3 | 93.7% | 90.2% | – |
| Terminal-Bench 2.1 | 89.9% | 91.4% | 87.6% |
| DeepSWE | 74.1% | – | 73.8% |
| Humanity's Last Exam | 54.7% | 59.1% | 47.8% |
| Artificial Analysis Coding Index | 77.1% | 81.6% | 76.3% |
| Artificial Analysis Intelligence Index | 54.7% | 56.8% | 47.1% |

Terminal-Bench 2.1: Fable 5.1 Leads, Astra Edges Gemini

Claude Fable 5.1 leads Terminal-Bench 2.1 at 91.4% at max, with GPT-6 Astra second at 89.9% at high and Gemini 3.8 Flash seventh at 87.6% at high 1. Chubby♨️ declared that Gemini 3.8 Flash outperforms GPT-5.6 Sol and Claude Opus 5 on Terminal-Bench 2.1, but the scores show the opposite: GPT-5.6 Sol at 89.5% and Claude Opus 5 at 89.1% both rank above Gemini 3.8 Flash at 87.6% 3.

Terminal-Bench 2.1: top 6 API-access and open-weight modelsTerminal-Bench 2.1: Claude Fable 5.1 (max) 91.4%, GPT-6 Astra (high) 89.9%, GPT-5.6 Sol (xhigh) 89.5% and others.Terminal-Bench 2.1: top 6 API-access and open-weight modelsAPI accessOpen weights0%25%50%75%100%Claude Fable 5.1 (max)Claude Fable 5.1 (max) · Anthropic · 91.4% · rank 191.4%GPT-6 Astra (high)GPT-6 Astra (high) · OpenAI · 89.9% · rank 289.9%GPT-5.6 Sol (xhigh)GPT-5.6 Sol (xhigh) · OpenAI · 89.5% · rank 389.5%Claude Opus 5 (max)Claude Opus 5 (max) · Anthropic · 89.1% · rank 489.1%Grok 4.6 (high)Grok 4.6 (high) · xAI · 88.4% · rank 588.4%GPT-5.6 Terra (max)GPT-5.6 Terra (max) · OpenAI · 88.0% · rank 688.0%Gemini 3.8 Flash (high)Gemini 3.8 Flash (high) · Google DeepMind · 87.6% · rank 787.6%Kimi K3 (max)Kimi K3 (max) · Moonshot · 85.0% · rank 1185.0%GLM-5.3-FlashGLM-5.3-Flash · Z.ai (Zhipu AI) · 84.3% · rank 1484.3%DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max) · DeepSeek · 78.7% · rank 2578.7%DeepSeek V4 Pro 0813 (max)DeepSeek V4 Pro 0813 (max) · DeepSeek · 78.7% · rank 2678.7%GLM-5.2 (max)GLM-5.2 (max) · Z.ai (Zhipu AI) · 77.9% · rank 2977.9%Kimi K2.7 CodeKimi K2.7 Code · Moonshot · 67.4% · rank 4467.4%Source: Artificial Analysis, https://artificialanalysis.ai/ Licence Free API, attribution required. Retrieved 2026-09-05.
Terminal-Bench 2.1: Claude Fable 5.1 (max) 91.4%, GPT-6 Astra (high) 89.9%, GPT-5.6 Sol (xhigh) 89.5% and others.

Data table

| Rank | Model | Access | Score |
| ---: | --- | --- | ---: |
| 1 | Claude Fable 5.1 (max) | API | 91.4% |
| 2 | GPT-6 Astra (high) | API | 89.9% |
| 3 | GPT-5.6 Sol (xhigh) | API | 89.5% |
| 4 | Claude Opus 5 (max) | API | 89.1% |
| 5 | Grok 4.6 (high) | API | 88.4% |
| 6 | GPT-5.6 Terra (max) | API | 88.0% |
| 7 | Gemini 3.8 Flash (high) | API | 87.6% |
| 11 | Kimi K3 (max) | Open weights | 85.0% |
| 14 | GLM-5.3-Flash | Open weights | 84.3% |
| 25 | DeepSeek V4 Flash 0731 (max) | Open weights | 78.7% |
| 26 | DeepSeek V4 Pro 0813 (max) | Open weights | 78.7% |
| 29 | GLM-5.2 (max) | Open weights | 77.9% |
| 44 | Kimi K2.7 Code | Open weights | 67.4% |

DeepSWE: Astra Leads, Gemini Close Behind at Lower Cost

On DeepSWE, GPT-6 Astra leads with 74.1% at xhigh, Gemini 3.8 Flash follows at 73.8% at high, and Claude Fable 5.1 has no DeepSWE score 1. The cost picture matters here: GPT-6 Astra spent $6.52 per task on average, Gemini 3.8 Flash $2.36. For teams weighing software-engineering agents by what a task costs, Flash delivers nearly the same score at roughly a third of the spend.

DeepSWE against cost per task28 models plotted by DeepSWE score against cost per task.DeepSWE against cost per taskAPI accessOpen weights0%25%50%75%100%$0.10$0.30$1.00$3.00$10$30mean cost per task in USD, log scaleClaude Fable 5 · 69.9% · $13/1MClaude Opus 4.8 · 59.0% · $13/1MClaude Opus 5 · 73.6% · $12/1MClaude Sonnet 4.6 · 29.9% · $5.52/1MClaude Sonnet 5 · 53.8% · $26/1MDeepSeek-V4-Flash · 53.3% · $0.10/1MDeepSeek-V4-Pro · 62.8% · $0.24/1MGemini 3.1 Pro Preview · 11.7% · $2.14/1MGemini 3.5 Flash · 36.1% · $3.45/1MGemini 3.6 Flash · 46.7% · $4.42/1MGemini 3.7 Flash · 65.5% · $2.03/1MGemini 3.8 Flash · 73.8% · $2.36/1MGLM-5.2 · 43.8% · $3.92/1MGLM-5.3-Flash · 63.4% · $0.48/1MGLM-5.3 · 69.0% · $3.99/1MGPT-5.4 · 51.8% · $5.65/1MGPT-5.5 · 67.0% · $7.23/1MGPT-5.6 Luna · 67.2% · $3.03/1MGPT-5.6 Sol · 72.7% · $8.39/1MGPT-5.6 Terra · 69.6% · $4.95/1MGPT-6 Astra · 74.1% · $6.52/1MGrok 4.5 · 53.8% · $2.42/1MGrok 4.6 · 67.5% · $3.45/1MKimi K2.7 Code · 30.5% · $2.82/1MKimi K3 · 68.5% · $4.65/1MMuse Spark 1.1 · 53.3% · $2.36/1MMuse Spark 1.2 · 54.9% · $3.70/1MQwen3.8 Max · 57.5% · $3.73/1MGemini 3.8 FlashGPT-6 AstraClaude Opus 5GPT-5.6 TerraDeepSeek-V4-FlashDeepSeek-V4-ProSource: Datacurve, DeepSWE leaderboard v1.1, https://deepswe.datacurve.ai/. Licence not stated (public leaderboard, cited with attribution).Retrieved 2026-09-05.Cost is the mean spend per task over four runs on mini-swe-agent, as measured by Datacurve.
28 models plotted by DeepSWE score against cost per task.

Data table

| Model | Access | Score | $/task |
| --- | --- | ---: | ---: |
| GPT-6 Astra | API | 74.1% | $6.52 |
| Gemini 3.8 Flash | API | 73.8% | $2.36 |
| Claude Opus 5 | API | 73.6% | $12 |
| GPT-5.6 Sol | API | 72.7% | $8.39 |
| Claude Fable 5 | API | 69.9% | $13 |
| GPT-5.6 Terra | API | 69.6% | $4.95 |
| GLM-5.3 | API | 69.0% | $3.99 |
| Kimi K3 | Open weights | 68.5% | $4.65 |
| Grok 4.6 | API | 67.5% | $3.45 |
| GPT-5.6 Luna | API | 67.2% | $3.03 |
| GPT-5.5 | API | 67.0% | $7.23 |
| Gemini 3.7 Flash | API | 65.5% | $2.03 |
| GLM-5.3-Flash | Open weights | 63.4% | $0.48 |
| DeepSeek-V4-Pro | Open weights | 62.8% | $0.24 |
| Claude Opus 4.8 | API | 59.0% | $13 |
| Qwen3.8 Max | API | 57.5% | $3.73 |
| Muse Spark 1.2 | API | 54.9% | $3.70 |
| Claude Sonnet 5 | API | 53.8% | $26 |
| Grok 4.5 | API | 53.8% | $2.42 |
| DeepSeek-V4-Flash | Open weights | 53.3% | $0.10 |
| Muse Spark 1.1 | API | 53.3% | $2.36 |
| GPT-5.4 | API | 51.8% | $5.65 |
| Gemini 3.6 Flash | API | 46.7% | $4.42 |
| GLM-5.2 | Open weights | 43.8% | $3.92 |
| Gemini 3.5 Flash | API | 36.1% | $3.45 |
| Kimi K2.7 Code | Open weights | 30.5% | $2.82 |
| Claude Sonnet 4.6 | API | 29.9% | $5.52 |
| Gemini 3.1 Pro Preview | API | 11.7% | $2.14 |

Humanity's Last Exam: Fable 5.1 on Top

Claude Fable 5.1 leads Humanity's Last Exam at 59.1% at max, ahead of GPT-6 Astra at 54.7% at max, with Gemini 3.8 Flash at 47.8% at high 1. Chubby♨️ noted Gemini 3.8 Flash outperforms GPT-5.6 Sol and Claude Opus 5 on HLE, but the scores contradict that claim: GPT-5.6 Sol at 49.5% and Claude Opus 5 at 54.9% both rank above Gemini 3.8 Flash at 47.8% 3.

Artificial Analysis Coding Index: Fable 5.1 Leads

Claude Fable 5.1 tops the Artificial Analysis Coding Index at 81.6% at max, while GPT-6 Astra sits at 77.1% at high and Gemini 3.8 Flash at 76.3% at high 1. Reddy's assertion that Fable 5.1 is the king of coding matches the ranking here 8. Simon Willison's hands-on experience with Fable 5.1 producing the best SVG pelican from any Anthropic model, at a cost of $3.30, reflects the max-effort setting that drives these scores 5.

Artificial Analysis Intelligence Index: Fable 5.1 Leads Astra

The Artificial Analysis Intelligence Index, a composite measure, places Claude Fable 5.1 first at 56.8% at max, GPT-6 Astra second at 54.7% at max, and Gemini 3.8 Flash seventh at 47.1% at high 1. Reddy claimed Astra beats Fable 5.1 on reasoning, but the Intelligence Index ranking shows Fable 5.1 ahead of Astra 8.

Price and Speed Context

GPT-6 Astra costs $20.00 per 1M tokens blended with no published speed 2. Claude Fable 5.1 costs $20.00 per 1M tokens blended at 71 tokens/s. Gemini 3.8 Flash costs $1.50 per 1M tokens blended at 429 tokens/s. All three are API-access models; none are open-weight.

Terminal-Bench 2.1 against price per million tokens137 models plotted by Terminal-Bench 2.1 score against price per million tokens.Terminal-Bench 2.1 against price per million tokensAPI accessOpen weights0%25%50%75%100%$0.10$0.30$1.00$3.00$10$30price per 1M tokens, log scaleAgnes 2.5 Pro Alpha · 67.0% · $0.56/1MAgnes 2.5 Pro Beta · 69.7% · $0.15/1MApodex 1.1 · 69.7% · $0.97/1MCeleris-1 · 11.2% · $0.33/1MClaude Fable 5.1 · 91.4% · $20/1MClaude Fable 5 · 84.6% · $20/1MClaude Haiku 4.5 · 44.2% · $2.00/1MClaude Opus 4.7 · 83.1% · $10/1MClaude Opus 4.8 · 84.6% · $10/1MClaude Opus 5 · 89.1% · $10/1MClaude Sonnet 4.5 · 55.8% · $6.00/1MClaude Sonnet 4.6 · 71.2% · $6.00/1MClaude Sonnet 4 · 36.3% · $6.00/1MClaude Sonnet 5 · 80.5% · $4.00/1MCommand A+ · 22.8% · $4.38/1MDeepSeek R1 (Jan '25) · 19.1% · $2.50/1MDeepSeek V3 0324 · 13.9% · $0.48/1MDeepSeek-V3.1-Terminus · 44.9% · $0.45/1MDeepSeek-V3.2 · 46.8% · $0.32/1MDeepSeek V3 (Dec '24) · 16.9% · $0.49/1MDeepSeek V4 Flash 0731 · 78.7% · $0.66/1MDeepSeek V4 Flash Vision · 74.2% · $0.66/1MDeepSeek-V4-Flash · 61.8% · $0.12/1MDeepSeek V4 Pro 0813 · 78.7% · $1.98/1MDeepSeek-V4-Pro · 64.8% · $0.54/1MGemini 2.5 Pro · 28.5% · $3.44/1MGemini 3.1 Flash-Lite · 31.1% · $0.56/1MGemini 3.5 Flash-Lite · 53.6% · $0.85/1MGemini 3.5 Flash · 78.7% · $3.38/1MGemini 3.6 Flash · 77.5% · $1.50/1MGemini 3.7 Flash · 85.8% · $1.50/1MGemini 3.8 Flash · 87.6% · $1.50/1MGemma 4 12B · 27.3% · $0.15/1MGemma 4 26B A4B · 39.0% · $0.18/1MGemma 4 31B · 43.4% · $0.20/1MGemma 4 E4B · 1.9% · $0.04/1MGLM-4.6 · 49.4% · $0.96/1MGLM-4.7 · 45.3% · $1.00/1MGLM-5.1 · 61.8% · $2.00/1MGLM-5.2 · 77.9% · $2.15/1MGLM-5.3-Flash · 84.3% · $0.24/1MGLM-5.3 · 83.9% · $2.15/1MGPT-4.1 mini · 10.1% · $0.70/1MGPT-4.1 nano · 3.7% · $0.17/1MGPT-4o mini · 5.6% · $0.26/1MGPT-5.1 · 52.4% · $3.44/1MGPT-5.4 Mini · 59.2% · $1.69/1MGPT-5.4 Nano · 60.7% · $0.46/1MGPT-5.4 · 78.3% · $5.63/1MGPT-5.5 Instant (June 2026) · 34.8% · $11/1MGPT-5.5 · 84.3% · $11/1MGPT-5.6 Luna · 80.9% · $0.45/1MGPT-5.6 Sol · 89.5% · $8.00/1MGPT-5.6 Terra · 88.0% · $4.50/1MGPT-5 mini · 3.7% · $0.69/1MGPT-5 · 35.2% · $3.44/1MGPT-6 Astra · 89.9% · $20/1Mgpt-oss-120b · 26.2% · $0.26/1Mgpt-oss-20b · 13.9% · $0.09/1MGranite 4.1 8B · 3.4% · $0.06/1MGranite 4.2 30B · 26.6% · $0.28/1MGranite 4.2 3B · 13.9% · $0.05/1MGranite 4.2 8B · 18.4% · $0.11/1MGrok 4.3 · 39.7% · $1.56/1MGrok 4.5 · 81.6% · $3.00/1MGrok 4.6 · 88.4% · $3.00/1MGrok Build 0.1 0616 · 52.1% · $1.25/1MHy3 · 64.4% · $0.24/1MHyperNova 60B 2605 (high, based on gpt-oss-120b) · 18.4% · $0.07/1MInkling-Small · 55.1% · $0.53/1MInkling · 55.1% · $1.76/1MKAT Coder Pro V2 · 70.0% · $0.53/1MKimi K2.5 · 45.7% · $1.14/1MKimi K2.6 · 65.9% · $1.71/1MKimi K2.7 Code · 67.4% · $1.71/1MKimi K3 · 85.0% · $6.00/1MLing 3.0 Flash · 55.4% · $0.11/1MLlama 3.1 Instruct 8B · 1.5% · $0.03/1MLlama 3.3 Instruct 70B · 4.9% · $0.67/1MLlama 4 Maverick · 7.9% · $0.42/1MLlama 4 Scout · 3.7% · $0.31/1MLongCat 2.0 · 50.2% · $1.30/1MMagistral Small 1.2 · 4.5% · $0.75/1MMercury 2 · 27.3% · $0.38/1MMiMo-V2.5-Pro · 65.2% · $0.54/1MMiMo-V2.5 · 63.7% · $0.17/1MMiMo-V2-Flash · 61.8% · $0.15/1MMiniMax-M2.7 · 55.4% · $0.53/1MMiniMax-M3 · 65.2% · $0.53/1MMinistral 3 14B · 9.7% · $0.20/1MMinistral 3 3B · 0.0% · $0.10/1MMinistral 3 8B · 4.1% · $0.15/1MMistral Large 3 · 12.0% · $0.75/1MMistral Medium 3.1 · 13.9% · $0.80/1MMistral Medium 3.5 · 50.6% · $3.00/1MMistral Small 3.1 · 26.2% · $0.15/1MMistral Small 3.2 · 5.6% · $0.15/1MMistral Small 4 · 21.0% · $0.26/1MMuse Glimmer · 51.7% · $0.64/1MMuse Spark 1.1 · 77.9% · $2.00/1MMuse Spark 1.2 · 80.1% · $2.00/1MMuse Spark 1.3 · 85.8% · $2.00/1MNemotron 3.5 Lightning · 24.3% · $0.10/1MNemotron 3 Nano Omni 30B A3B Reasoning · 6.7% · $0.13/1MNemotron 3 Super 120B A12B · 38.6% · $0.35/1MNemotron 3 Ultra 550B A55B · 53.9% · $1.10/1MNex-N2-Pro · 67.8% · $1.00/1MNova 2.0 Lite · 16.1% · $0.85/1MNova 2.0 Pro Preview · 29.6% · $3.44/1MNVIDIA Nemotron 3 Nano 30B A3B · 6.7% · $0.09/1Mo3-mini · 4.5% · $1.93/1MQuasar 438B (max, based on GLM-5.2) · 69.3% · $0.90/1MQwen3-14B · 4.9% · $1.31/1MQwen3 235B A22B 2507 · 12.0% · $0.75/1MQwen3 30B A3B 2507 · 1.5% · $0.75/1MQwen3-32B · 5.2% · $0.28/1MQwen3.5-122B-A10B · 47.6% · $1.10/1MQwen3.5-35B-A3B · 40.8% · $0.69/1MQwen3.5 397B-A17B · 51.3% · $1.35/1MQwen3.5-4B · 25.8% · $0.06/1MQwen3.5-9B · 29.2% · $0.15/1MQwen3.6 27B · 60.7% · $1.35/1MQwen3.6 35B A3B · 44.9% · $0.84/1MQwen3.6 Plus · 61.4% · $1.13/1MQwen3.7-Max · 74.5% · $3.75/1MQwen3.7-Plus · 61.0% · $0.70/1MQwen3.8 2.4T A95B · 82.0% · $3.00/1MQwen3.8 27B · 79.8% · $1.13/1MQwen3.8-Flash-Next · 86.1% · $0.23/1MQwen3-8B · 2.2% · $0.31/1Mqwen3-coder-next · 38.2% · $0.56/1MQwen3 Next 80B A3B · 6.7% · $0.41/1MRing-2.6-1T · 43.1% · $0.85/1MSolar Pro 3 · 12.0% · $0.26/1MSolar Pro 4 · 57.3% · $0.53/1MStep 3.7 Flash · 39.3% · $0.44/1MTrinity Large Thinking · 20.6% · $0.41/1MClaude Fable 5.1Gemini 3.8 FlashGPT-6 AstraLlama 3.1 Instruct 8BSource: Artificial Analysis, https://artificialanalysis.ai/ Licence Free API, attribution required. Retrieved 2026-09-05.Price is the blended cost per million tokens at 3:1 input to output, from Artificial Analysis (https://artificialanalysis.ai/), attributionrequired.
137 models plotted by Terminal-Bench 2.1 score against price per million tokens.

38 scored models without a published price are not plotted.

Data table

| Model | Access | Score | $/1M tokens |
| --- | --- | ---: | ---: |
| Claude Fable 5.1 | API | 91.4% | $20 |
| GPT-6 Astra | API | 89.9% | $20 |
| GPT-5.6 Sol | API | 89.5% | $8.00 |
| Claude Opus 5 | API | 89.1% | $10 |
| Grok 4.6 | API | 88.4% | $3.00 |
| GPT-5.6 Terra | API | 88.0% | $4.50 |
| Gemini 3.8 Flash | API | 87.6% | $1.50 |
| Qwen3.8-Flash-Next | API | 86.1% | $0.23 |
| Gemini 3.7 Flash | API | 85.8% | $1.50 |
| Muse Spark 1.3 | API | 85.8% | $2.00 |
| Kimi K3 | Open weights | 85.0% | $6.00 |
| Claude Fable 5 | API | 84.6% | $20 |
| Claude Opus 4.8 | API | 84.6% | $10 |
| GLM-5.3-Flash | Open weights | 84.3% | $0.24 |
| GPT-5.5 | API | 84.3% | $11 |
| GLM-5.3 | API | 83.9% | $2.15 |
| Claude Opus 4.7 | API | 83.1% | $10 |
| Qwen3.8 2.4T A95B | API | 82.0% | $3.00 |
| Grok 4.5 | API | 81.6% | $3.00 |
| GPT-5.6 Luna | API | 80.9% | $0.45 |
| Claude Sonnet 5 | API | 80.5% | $4.00 |
| Muse Spark 1.2 | API | 80.1% | $2.00 |
| Qwen3.8 27B | API | 79.8% | $1.13 |
| DeepSeek V4 Flash 0731 | Open weights | 78.7% | $0.66 |
| DeepSeek V4 Pro 0813 | Open weights | 78.7% | $1.98 |
| Gemini 3.5 Flash | API | 78.7% | $3.38 |
| GPT-5.4 | API | 78.3% | $5.63 |
| GLM-5.2 | Open weights | 77.9% | $2.15 |
| Muse Spark 1.1 | API | 77.9% | $2.00 |
| Gemini 3.6 Flash | API | 77.5% | $1.50 |
| Qwen3.7-Max | API | 74.5% | $3.75 |
| DeepSeek V4 Flash Vision | API | 74.2% | $0.66 |
| Claude Sonnet 4.6 | API | 71.2% | $6.00 |
| KAT Coder Pro V2 | API | 70.0% | $0.53 |
| Agnes 2.5 Pro Beta | API | 69.7% | $0.15 |
| Apodex 1.1 | API | 69.7% | $0.97 |
| Quasar 438B (max, based on GLM-5.2) | API | 69.3% | $0.90 |
| Nex-N2-Pro | API | 67.8% | $1.00 |
| Kimi K2.7 Code | Open weights | 67.4% | $1.71 |
| Agnes 2.5 Pro Alpha | API | 67.0% | $0.56 |
| Kimi K2.6 | Open weights | 65.9% | $1.71 |
| MiMo-V2.5-Pro | Open weights | 65.2% | $0.54 |
| MiniMax-M3 | Open weights | 65.2% | $0.53 |
| DeepSeek-V4-Pro | Open weights | 64.8% | $0.54 |
| Hy3 | API | 64.4% | $0.24 |
| MiMo-V2.5 | Open weights | 63.7% | $0.17 |
| DeepSeek-V4-Flash | Open weights | 61.8% | $0.12 |
| GLM-5.1 | Open weights | 61.8% | $2.00 |
| MiMo-V2-Flash | Open weights | 61.8% | $0.15 |
| Qwen3.6 Plus | API | 61.4% | $1.13 |
| Qwen3.7-Plus | API | 61.0% | $0.70 |
| GPT-5.4 Nano | API | 60.7% | $0.46 |
| Qwen3.6 27B | Open weights | 60.7% | $1.35 |
| GPT-5.4 Mini | API | 59.2% | $1.69 |
| Solar Pro 4 | API | 57.3% | $0.53 |
| Claude Sonnet 4.5 | API | 55.8% | $6.00 |
| Ling 3.0 Flash | API | 55.4% | $0.11 |
| MiniMax-M2.7 | Open weights | 55.4% | $0.53 |
| Inkling-Small | Open weights | 55.1% | $0.53 |
| Inkling | Open weights | 55.1% | $1.76 |
| Nemotron 3 Ultra 550B A55B | API | 53.9% | $1.10 |
| Gemini 3.5 Flash-Lite | API | 53.6% | $0.85 |
| GPT-5.1 | API | 52.4% | $3.44 |
| Grok Build 0.1 0616 | API | 52.1% | $1.25 |
| Muse Glimmer | API | 51.7% | $0.64 |
| Qwen3.5 397B-A17B | Open weights | 51.3% | $1.35 |
| Mistral Medium 3.5 | Open weights | 50.6% | $3.00 |
| LongCat 2.0 | API | 50.2% | $1.30 |
| GLM-4.6 | Open weights | 49.4% | $0.96 |
| Qwen3.5-122B-A10B | Open weights | 47.6% | $1.10 |
| DeepSeek-V3.2 | Open weights | 46.8% | $0.32 |
| Kimi K2.5 | Open weights | 45.7% | $1.14 |
| GLM-4.7 | Open weights | 45.3% | $1.00 |
| DeepSeek-V3.1-Terminus | Open weights | 44.9% | $0.45 |
| Qwen3.6 35B A3B | API | 44.9% | $0.84 |
| Claude Haiku 4.5 | API | 44.2% | $2.00 |
| Gemma 4 31B | API | 43.4% | $0.20 |
| Ring-2.6-1T | Open weights | 43.1% | $0.85 |
| Qwen3.5-35B-A3B | Open weights | 40.8% | $0.69 |
| Grok 4.3 | API | 39.7% | $1.56 |
| Step 3.7 Flash | Open weights | 39.3% | $0.44 |
| Gemma 4 26B A4B | Open weights | 39.0% | $0.18 |
| Nemotron 3 Super 120B A12B | API | 38.6% | $0.35 |
| qwen3-coder-next | API | 38.2% | $0.56 |
| Claude Sonnet 4 | API | 36.3% | $6.00 |
| GPT-5 | API | 35.2% | $3.44 |
| GPT-5.5 Instant (June 2026) | API | 34.8% | $11 |
| Gemini 3.1 Flash-Lite | API | 31.1% | $0.56 |
| Nova 2.0 Pro Preview | API | 29.6% | $3.44 |
| Qwen3.5-9B | Open weights | 29.2% | $0.15 |
| Gemini 2.5 Pro | API | 28.5% | $3.44 |
| Gemma 4 12B | API | 27.3% | $0.15 |
| Mercury 2 | API | 27.3% | $0.38 |
| Granite 4.2 30B | API | 26.6% | $0.28 |
| gpt-oss-120b | Open weights | 26.2% | $0.26 |
| Mistral Small 3.1 | Open weights | 26.2% | $0.15 |
| Qwen3.5-4B | Open weights | 25.8% | $0.06 |
| Nemotron 3.5 Lightning | API | 24.3% | $0.10 |
| Command A+ | Open weights | 22.8% | $4.38 |
| Mistral Small 4 | Open weights | 21.0% | $0.26 |
| Trinity Large Thinking | API | 20.6% | $0.41 |
| DeepSeek R1 (Jan '25) | API | 19.1% | $2.50 |
| Granite 4.2 8B | API | 18.4% | $0.11 |
| HyperNova 60B 2605 (high, based on gpt-oss-120b) | API | 18.4% | $0.07 |
| DeepSeek V3 (Dec '24) | API | 16.9% | $0.49 |
| Nova 2.0 Lite | API | 16.1% | $0.85 |
| DeepSeek V3 0324 | API | 13.9% | $0.48 |
| gpt-oss-20b | Open weights | 13.9% | $0.09 |
| Granite 4.2 3B | API | 13.9% | $0.05 |
| Mistral Medium 3.1 | API | 13.9% | $0.80 |
| Mistral Large 3 | Open weights | 12.0% | $0.75 |
| Qwen3 235B A22B 2507 | API | 12.0% | $0.75 |
| Solar Pro 3 | API | 12.0% | $0.26 |
| Celeris-1 | API | 11.2% | $0.33 |
| GPT-4.1 mini | API | 10.1% | $0.70 |
| Ministral 3 14B | API | 9.7% | $0.20 |
| Llama 4 Maverick | Open weights | 7.9% | $0.42 |
| Nemotron 3 Nano Omni 30B A3B Reasoning | API | 6.7% | $0.13 |
| NVIDIA Nemotron 3 Nano 30B A3B | API | 6.7% | $0.09 |
| Qwen3 Next 80B A3B | API | 6.7% | $0.41 |
| GPT-4o mini | API | 5.6% | $0.26 |
| Mistral Small 3.2 | Open weights | 5.6% | $0.15 |
| Qwen3-32B | Open weights | 5.2% | $0.28 |
| Llama 3.3 Instruct 70B | API | 4.9% | $0.67 |
| Qwen3-14B | Open weights | 4.9% | $1.31 |
| Magistral Small 1.2 | Open weights | 4.5% | $0.75 |
| o3-mini | API | 4.5% | $1.93 |
| Ministral 3 8B | API | 4.1% | $0.15 |
| GPT-4.1 nano | API | 3.7% | $0.17 |
| GPT-5 mini | API | 3.7% | $0.69 |
| Llama 4 Scout | Open weights | 3.7% | $0.31 |
| Granite 4.1 8B | API | 3.4% | $0.06 |
| Qwen3-8B | Open weights | 2.2% | $0.31 |
| Gemma 4 E4B | API | 1.9% | $0.04 |
| Llama 3.1 Instruct 8B | API | 1.5% | $0.03 |
| Qwen3 30B A3B 2507 | API | 1.5% | $0.75 |
| Ministral 3 3B | API | 0.0% | $0.10 |

Verdict by Use

For mathematical research, pick GPT-6 Astra — it leads both FrontierMath tiers, and its nearest competitor on Tier 4, Claude Fable 5.1, is nearly ten points behind at 87.8%. For agentic coding and terminal work, pick Claude Fable 5.1, provided you accept $20.00 per 1M tokens blended, since it leads Terminal-Bench 2.1, the Coding Index, and Humanity's Last Exam. For software engineering at scale on DeepSWE, pick Gemini 3.8 Flash if cost per task matters: it trails Astra on DeepSWE at $2.36 mean cost per task versus $6.52, and at $1.50 per 1M tokens blended it is the cheapest of the three. Ethan Mollick's demonstrations of Astra building a 3D Zork adaptation and procedural ocean simulations suggest it excels at complex creative coding tasks, but those are relayed anecdotes rather than benchmark scores 79.

AI Tools

    GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash: Benchmark Head-to-Head | TerraNet Technologies