ai-models
Claude Opus 5
Anthropic's recommended starting model for agentic coding and enterprise work, with Fable-class intelligence at half the price
Starts at
From $5/1M input tokens
Pricing tier: Usage-Based
Visit Claude Opus 5Independent software comparison
Cost and speed vs. frontier reasoning depth
ai-models · high search interest
ai-models
Anthropic's recommended starting model for agentic coding and enterprise work, with Fable-class intelligence at half the price
Starts at
From $5/1M input tokens
Pricing tier: Usage-Based
Visit Claude Opus 5ai-models
Anthropic's top-tier model for demanding reasoning and long-horizon agentic work, at $10 in and $50 out per million tokens
Starts at
From $10/1M input tokens
Pricing tier: Usage-Based
Visit Claude Fable 5.1Expert analysis
Deploying Anthropic's hosted models involves navigating a direct operational split between the economical throughput of Claude Opus 5 and the deeper reasoning depth of Claude Fable 5.1. Both systems provide a one-million-token context window, structured tool use, multimodal image input, and broad cloud availability across the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry. However, their execution characteristics and operational profiles diverge significantly. Claude Opus 5 pairs a baseline price of $5 per million input tokens and $25 per million output tokens with moderate latency, toggleable thinking, and an optional 2.5-times fast mode. Claude Fable 5.1 doubles standard rates to $10 in and $50 out, pairing that premium with Anthropic's deepest reasoning for complex research, spreadsheets, and long-horizon tasks, but constraining integration through slower speeds, mandatory thinking, and an error on forced tool choice.
Feature matrix
Rows are grouped by capability, and each cell shows the wording from that vendor’s own documentation. “Not documented” means we found no cited source for that capability, which is not the same as the product lacking it.
| Capability | Claude Opus 5 | Claude Fable 5.1 |
|---|---|---|
| Starting price | From $5/1M input tokens | From $10/1M input tokens |
| Free plan | No | No |
| API available | Product API available | Product API available |
| Context window and output limit | 1M-token context window, 128K-token output, 300K on the Batch API in beta | 1M-token context window, 128K-token output |
| Thinking mode and effort control | Adaptive thinking on by default, effort from low to max | Adaptive thinking always on, steered by effort |
| Tool use and agent support | Tool use, including mid-conversation tool changes | Tool use, without forced tool choice |
| Image and document input | Text and image input, text output | Text and image input, text output |
| Long-horizon and autonomous work | Largest gains on deep reasoning, agentic and long-horizon tasks | Long-running agentic coding and multistep research |
| Speed and latency | Moderate comparative latency, with a fast mode | Slower comparative latency |
| Caching, batch and speed pricing | Cache reads at 10% of input price, Batch API at 50% off, fast mode at double price | Cache reads at 2.5% of input price, Batch API at 50% off |
| Open weights and licence terms | No published weights; API access only | No published weights; API access only |
| Running it on your own hardware | Hosted only; no self-hosted deployment offered | Hosted only; no self-hosted deployment offered |
| Cloud platform availability | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Availability in the vendor's own apps | Claude.ai, Claude Code and Claude Cowork | Claude.ai, Claude Code and Claude Enterprise |
| Safeguards and refusal handling | Not documented | Classifier refusals with fallback to another model |
| Knowledge cutoff and retirement date | Released 24 July 2026; knowledge cutoff May 2026; retirement not sooner than 24 July 2027 | Released 1 September 2026; knowledge cutoff June 2026; retirement not sooner than 1 September 2027 |
Model benchmarks
Claude Opus 5 runs on Claude Opus 5 and Claude Fable 5.1 on Claude Fable 5.1. These are the models’ scores, not the tools’: independent evaluations from Epoch AI, Artificial Analysis and Datacurve, each at the model’s best published effort setting, last read 2026-09-28. A dash means the model has not been scored on that benchmark yet.
| Benchmark | Claude Opus 5 | Claude Fable 5.1 |
|---|---|---|
| FrontierMath Tier 4 | 73.2% at max | 87.8% at max |
| FrontierMath Tiers 1–3 | 85.6% at max | 90.2% at max |
| ARC-AGI-2 | 90.4% † at max | 90.0% † at max |
| Terminal-Bench 2.1 | 89.1% at max | 91.4% at max |
| DeepSWE | 73.6% at max | – |
| Humanity's Last Exam | 54.9% at max | 59.1% at max |
| Artificial Analysis Coding Index | 78.0% at max | 81.6% at max |
Sources: Artificial Analysis · Epoch AI · Datacurve.
† Relayed by the source from a vendor or external leaderboard rather than run by it.
Detailed comparison
The two models implement adaptive thinking under distinct operational parameters that dictate runtime latency and programmatic control. Claude Opus 5 operates with adaptive thinking enabled by default at a high effort setting, but it allows developers to configure effort from low to max or disable internal thinking entirely at effort high or below for latency-critical pipelines. In Anthropic's comparative lineup, Opus 5 exhibits moderate latency and offers an exclusive fast mode on the Claude API that increases processing speed by roughly 2.5 times at $10 in and $50 out. In contrast, Claude Fable 5.1 treats adaptive thinking as mandatory. Thinking is permanently active and cannot be toggled off, though developers can steer its reasoning depth using effort parameters or per-message effort adjustments in beta. Because Fable 5.1 spends compute verifying and iterating through complex reasoning paths, Anthropic rates its comparative latency as Slower in the current lineup. Teams adopting Fable 5.1 must design asynchronous, background workflows that accommodate these longer turnaround times.
Migrating autonomous agents between Opus 5 and Fable 5.1 reveals key differences in API contracts and tool execution behaviors. Claude Opus 5 accommodates dynamic agent workflows by supporting mid-conversation tool adjustments, allowing an orchestration system to redefine available tool schemas across ongoing dialogue turns. It also supports up to 300,000 output tokens on the Batch API under a beta header. Claude Fable 5.1 introduces stricter architectural boundaries. A breaking change in Fable 5.1 is that forced tool choice returns an explicit API error; the model must evaluate and select its own tool calls via its internal reasoning loop rather than being programmatically coerced into executing a pre-assigned function. Furthermore, thinking blocks in Fable 5.1 are bound to the model that produced them, preventing architectures from replaying chain-of-thought blocks across different models. While both systems handle autonomous coding and multistep execution, existing agent harnesses that rely on deterministic, forced function calls require structural code revisions before integrating Fable 5.1.
The commercial structure of these models presents a distinct cost-allocation trade-off that depends heavily on prompt caching efficiency. At standard input and output rates, Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, exactly half the list price of Claude Fable 5.1, which stands at $10 per million input tokens and $50 per million output tokens. Asynchronous execution through the Batch API cuts rates in half for both models, bringing Opus 5 down to $2.50 in and $12.50 out, while Fable 5.1 drops to $5 in and $25 out. For low-latency workloads, Opus 5 provides a fast mode at $10 in and $50 out on the Claude API. However, the economic balance shifts when dealing with massive, repetitive contexts. Fable 5.1 discounts prompt cache reads down to 2.5 percent of its base input price, charging $0.25 per million tokens compared to the standard 10 percent discount on Opus 5, which prices cache reads at $0.50 per million tokens. Long-running agent workflows that repeatedly query static codebases or large document caches see narrower cost disparities than list pricing suggests.
Enterprise implementations must navigate differing operational behaviors regarding safety classifiers and execution boundaries. Claude Fable 5.1 functions as Anthropic's deepest reasoning system for demanding workloads like complex spreadsheets, slides, and multi-step research, but its autonomous depth comes with distinct guardrail mechanics. When safeguards intervene during automated benchmark tasks, Fable 5.1 scores zero due to classifier refusals, and alignment evaluations show it can still sometimes bypass approvals and auto-mode classifiers. To maintain continuity, Anthropic provides a formal refusals-and-fallback guide specifically for Fable 5.1, instructing engineering teams to detect automated classifier refusals and programmatically retry stalled tasks on another model. In contrast, Claude Opus 5 operates with predictable moderation across Claude.ai, Claude Code, and Claude Cowork. Its lower default latency, flexible thinking toggles, and retirement commitment extending at least through July 2027 make it an operationally straightforward target for production software pipelines.
Best use case for Claude Opus 5
Claude Opus 5 suits engineering teams seeking an economical, versatile default model for agentic coding and general enterprise tasks with optional fast-mode execution.
Best use case for Claude Fable 5.1
Claude Fable 5.1 suits developers running complex, multistep research or demanding reasoning workloads where Opus 5 evaluations fall short and aggressive prompt cache discounts offset premium pricing.
Decision framework
Choose Claude Opus 5 when your operational priority centers on economical baseline token pricing, moderate latency, and flexible programmatic controls. It is the practical fit for high-volume agentic coding, general enterprise workflows, and architectures that require the option to disable thinking or accelerate execution through API fast mode. Move to Claude Fable 5.1 when internal evaluations on Opus 5 fall short on complex, multistep research, demanding mathematical reasoning, or dense document, spreadsheet, and slide operations. Fable 5.1 is the appropriate choice when tasks require peak reasoning depth, the engineering architecture can handle slower latency and mandatory thinking without forced tool choice, and frequent prompt cache hits at $0.25 per million tokens help offset the doubled list rates.
Bottom line
Claude Opus 5 costs half as much as Claude Fable 5.1 on standard input and output tokens while delivering faster response times and flexible reasoning toggles, whereas Claude Fable 5.1 delivers Anthropic's deepest reasoning capabilities alongside slower comparative latency and strict programmatic constraints. For standard agentic engineering and general enterprise workloads, Opus 5 provides the more balanced operational foundation due to its $5 and $25 token rates, optional 2.5-times fast mode, and ability to disable thinking when latency matters. Claude Fable 5.1 belongs in pipelines where evaluations demonstrate that Opus 5 cannot resolve the underlying reasoning problem, provided the engineering stack can accommodate double the token expense, mandatory adaptive thinking, and an API contract that prohibits forced tool selection.
Sources and verification
The product facts have been checked against the sources below. The AI-assisted analysis was audited against these exact evidence records and approved by a human editor.
Last verified September 21, 2026
Last verified September 21, 2026
Editorial validation
Human-approvedApproved September 21, 2026 after an automated evidence audit using gemini-3.6-flash.
Read our comparison methodology and editorial policy, learn about TerraNet, or report a correction.
Common questions
Adaptive thinking can be disabled on Claude Opus 5 when the effort parameter is set to high or below, though it is active by default. On Claude Fable 5.1, adaptive thinking is permanently enabled and cannot be turned off under any configuration.
Claude Opus 5 prices prompt cache reads at $0.50 per million tokens, representing the standard 90 percent discount off its $5 input price. Claude Fable 5.1 offers a steeper 97.5 percent discount, pricing prompt cache reads at $0.25 per million tokens despite its higher $10 baseline input rate.
No. In Claude Fable 5.1, attempting to force a tool choice returns an API error. The model uses its internal adaptive thinking loop to determine when and how to call external tools.
Anthropic rates the comparative latency of Claude Opus 5 as Moderate, whereas Claude Fable 5.1 is rated as Slower. In addition, Claude Opus 5 supports an optional fast mode on the Claude API that runs at approximately 2.5 times default speed for $10 input and $50 output per million tokens.
AI-assisted draft audited against the cited product evidence and approved by a human editor. Vendor pricing and capabilities can change after the recorded verification date.
Continue researching
Claude Fable 5.1's always-on adaptive thinking, locked tool-choice mechanics, and premium rate of ten dollars per million input tokens and fifty dollars per million output tokens position it as a specialized engine for long-horizon agentic execution. Anthropic's own documentation explicitly directs engineering teams to begin with Claude Opus 5 for standard workloads, reserving Fable 5.1 primarily for scenarios where Opus evaluations at high effort levels still prove insufficient. When building high-throughput production systems, developers often encounter operational friction with Fable 5.1's slower latency profile, breaking changes such as returning an error upon forced tool selection, and the inability to deactivate reasoning tokens on straightforward tasks. Furthermore, organizations requiring dedicated self-hosting options, custom local deployments, or more permissive licensing frameworks cannot achieve those goals within Anthropic's hosted-only managed endpoints. Examining alternative hosted frontier systems and open-weight architectures allows development teams to calibrate their infrastructure specifically around latency requirements, input pricing, and deterministic runtime control.
Read guideClaude Opus 5 pairs a 1M-token context window and 128K-token output ceiling with adaptive thinking enabled by default, serving as Anthropic's recommended starting model for agentic coding and deep reasoning at $5 per million input tokens and $25 per million output tokens. However, its Moderate comparative latency rating positions it behind faster options in latency-sensitive pipelines, and its breaking changes—which keep thinking enabled unless manually turned down at effort high or below—can disrupt production configurations carried over from older versions. Furthermore, its research-preview fast mode doubles the token rates and remains restricted to the Claude API rather than third-party cloud environments, while organizations with sovereign infrastructure requirements cannot self-host its closed weights. These technical constraints, pricing structures, and runtime realities lead engineering teams to explore alternatives across Anthropic's portfolio, hyperscaler competitors, and open-weight architectures.
Read guideClaude Haiku 4.5 provides low-latency execution and high-volume cost efficiency at $1 per million input tokens, while Claude Sonnet 5 provides a 1M-token context window and autonomous tool planning at double the base token price. Teams optimizing for interactive user experiences, live customer support desks, and narrow margin footprints will find Haiku 4.5 the more practical fit. Conversely, projects requiring broad document synthesis, deep programmatic refactoring, and independent multi-turn agent loops will find Sonnet 5 essential despite its higher token counts and strict 400-error validation on sampling overrides.
Read guideClaude Sonnet 5 gives you a fast, cost-efficient workhorse priced at $2 in and $10 out per million tokens for everyday production throughput, whereas Claude Opus 5 gives you a frontier reasoning engine priced at $5 in and $25 out per million tokens built for complex autonomy and cybersecurity tasks. While both tools deploy identical 1M-token context buffers and cloud availability across the Claude API, AWS Bedrock, Google Cloud, and Microsoft Foundry, they should not be treated as interchangeable endpoints. Teams operating customer-facing interfaces, high-frequency tool pipelines, and latency-sensitive features will find Sonnet 5 far easier to sustain financially and operationally. Conversely, engineering departments deploying agents for multi-file refactoring, vulnerability inspection, and high-effort reasoning should absorb the cost of Opus 5, reserving Sonnet for the surrounding orchestration layers.
Read guideGLM-5.3-Flash delivers an accessible, lightweight 18B active-parameter architecture under an unencumbered MIT license with rock-bottom operational costs, whereas Kimi K3 demands a massive multi-node 2.8T-parameter footprint, custom legal approval, and always-on reasoning tokens to unlock frontier-level agentic task completion. For high-volume production, multimodal file pipelines, and internal deployments on standard server setups, GLM-5.3-Flash provides the most practical and legally clear path forward. For demanding terminal control, automated web browsing, and multi-step programmatic problem solving where accuracy supersedes operational cost, Kimi K3 stands as the superior agentic reasoning tool.
Read guideGPT-5.6 Sol gives developers a high-throughput endpoint with switchable reasoning and low base token prices, whereas Claude Fable 5.1 provides an agentic reasoning engine with mandatory thinking and deeply discounted prompt caching across multi-cloud infrastructure. Organizations managing fast, high-volume production queues will find Sol's $4.00 entry rate and zero-effort option ideal for maintaining strict budget and latency targets. Conversely, enterprise teams managing complex analytical problems across recurring reference contexts will gain superior stability and sustainable long-term economics from Fable 5.1.
Read guide