ai-models
Claude Sonnet 5
The Sonnet tier: 1M context, adaptive thinking and agentic ability close to the Opus tier at $2 in and $10 out per million tokens
Starts at
From $2/1M input tokens
Pricing tier: Usage-Based
Visit Claude Sonnet 5Independent software comparison
Cost and latency vs. deep reasoning and cybersecurity capability
ai-models · high search interest
ai-models
The Sonnet tier: 1M context, adaptive thinking and agentic ability close to the Opus tier at $2 in and $10 out per million tokens
Starts at
From $2/1M input tokens
Pricing tier: Usage-Based
Visit Claude Sonnet 5ai-models
Anthropic's recommended starting model for agentic coding and enterprise work, with Fable-class intelligence at half the price
Starts at
From $5/1M input tokens
Pricing tier: Usage-Based
Visit Claude Opus 5Expert analysis
Engineering leads and system architects integrating Anthropic models face a direct operational evaluation between Claude Sonnet 5 and Claude Opus 5. Both models share identical core boundaries: a 1M-token context window, 128K-token output limits on standard synchronous calls, and multimodal text and image comprehension. However, Anthropic positions them for sharply different runtime profiles. Claude Sonnet 5 is built for high-throughput, cost-sensitive production pipelines where low latency and autonomous tool operation matter most. Claude Opus 5 acts as Anthropic recommended starting model for deep reasoning, cybersecurity workloads, and multi-turn iterative autonomy, but commands two-and-a-half times the token cost alongside a slower base latency profile.
Feature matrix
Rows are grouped by capability, and each cell shows the wording from that vendor’s own documentation. “Not documented” means we found no cited source for that capability, which is not the same as the product lacking it.
| Capability | Claude Sonnet 5 | Claude Opus 5 |
|---|---|---|
| Starting price | From $2/1M input tokens | From $5/1M input tokens |
| Free plan | No | No |
| API available | Product API available | Product API available |
| Context window and output limit | 1M-token context window, 128K-token output, 300K on the Batch API in beta | 1M-token context window, 128K-token output, 300K on the Batch API in beta |
| Thinking mode and effort control | Adaptive thinking on by default, effort from low to max | Adaptive thinking on by default, effort from low to max |
| Tool use and agent support | Plans, uses browsers and terminals, and runs autonomously | Tool use, including mid-conversation tool changes |
| Image and document input | Text and image input, text output | Text and image input, text output |
| Long-horizon and autonomous work | Anthropic's most agentic Sonnet model yet | Largest gains on deep reasoning, agentic and long-horizon tasks |
| Speed and latency | Fast comparative latency | Moderate comparative latency, with a fast mode |
| Caching, batch and speed pricing | Cache reads at 10% of input price, Batch API at 50% off | Cache reads at 10% of input price, Batch API at 50% off, fast mode at double price |
| Open weights and licence terms | No published weights; API access only | No published weights; API access only |
| Running it on your own hardware | Hosted only; no self-hosted deployment offered | Hosted only; no self-hosted deployment offered |
| Cloud platform availability | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Availability in the vendor's own apps | Default model on Claude Free and Pro; available on Max, Team and Enterprise | Claude.ai, Claude Code and Claude Cowork |
| Safeguards and refusal handling | Cyber safeguards enabled by default | Not documented |
| Knowledge cutoff and retirement date | Released 30 June 2026; knowledge cutoff January 2026; retirement not sooner than 30 June 2027 | Released 24 July 2026; knowledge cutoff May 2026; retirement not sooner than 24 July 2027 |
Model benchmarks
Claude Sonnet 5 runs on Claude Sonnet 5 and Claude Opus 5 on Claude Opus 5. These are the models’ scores, not the tools’: independent evaluations from Epoch AI, Artificial Analysis and Datacurve, each at the model’s best published effort setting, last read 2026-09-28. A dash means the model has not been scored on that benchmark yet.
| Benchmark | Claude Sonnet 5 | Claude Opus 5 |
|---|---|---|
| FrontierMath Tier 4 | 29.3% at max | 73.2% at max |
| FrontierMath Tiers 1–3 | 65.6% at max | 85.6% at max |
| ARC-AGI-2 | – | 90.4% † at max |
| Terminal-Bench 2.1 | 80.5% at max | 89.1% at max |
| DeepSWE | 53.8% at max | 73.6% at max |
| Humanity's Last Exam | 41.3% at max | 54.9% at max |
| Artificial Analysis Coding Index | 71.5% at max | 78.0% at max |
Sources: Artificial Analysis · Epoch AI · Datacurve.
† Relayed by the source from a vendor or external leaderboard rather than run by it.
Detailed comparison
The direct expense of processing large contexts separates these two options immediately. Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens on standard API requests, an introductory rate that Anthropic made permanent. By contrast, Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens. For high-volume applications or continuous agentic loops, that price gap scales rapidly. Both models support prompt caching, where cache read hits drop to 10 percent of the base input token price ($0.20 per million for Sonnet 5 versus $0.50 per million for Opus 5). Cache writes cost $2.50 for five-minute retention and $4 for one-hour retention on Sonnet 5, compared to $6.25 and $10 on Opus 5. Both options cut rates by 50 percent when routing offline jobs through the Message Batches API, where Sonnet drops to $1 in and $5 out, and Opus drops to $2.50 in and $12.50 out. Teams should also remember that both models share a tokenizer that yields roughly 30 percent more tokens for identical text relative to Claude 4.6 architectures, making per-token cost comparisons across model generations non-linear.
Operational latency dictates which interactive interfaces each model can reliably support. In Anthropic comparative lineup rankings, Claude Sonnet 5 is rated as Fast, sitting just behind Haiku 4.5. Claude Opus 5 is classified as having Moderate comparative latency, reflecting the additional compute dedicated to its deeper reasoning paths. For teams deploying user-facing tools where multi-second turnarounds hurt engagement, Sonnet 5 provides a significantly snappier baseline. Opus 5 does introduce a dedicated research preview called Fast mode, available exclusively on the Claude API and not via partner clouds or batch endpoints. Fast mode runs at roughly 2.5 times the default Opus speed, but doubles unit rates to $10 per million input tokens and $50 per million output tokens. Unless a team has the budget to pay five times Sonnet rates to accelerate Opus reasoning, Sonnet 5 remains the natural selection for responsive conversational apps.
Both architectures rely on adaptive thinking enabled by default, with tunable effort settings ranging from low to max. However, the migration rules from earlier Claude versions reveal key technical constraints. On Claude Sonnet 5, manual extended thinking budgets and non-default sampling values (such as custom temperature, top_p, or top_k parameters) actively trigger 400 errors. Teams migrating legacy prompts tuned with manual token budgets or custom sampling must refactor their payload definitions to use the native effort parameter. On Claude Opus 5, adaptive thinking is similarly default with a base effort level of high, but disabling thinking entirely is restricted: requests can only turn off thinking when the effort level is set to high or below. Claude Opus 5 also supports mid-conversation tool modifications and excels at long-horizon task verification, allowing it to systematically test and iterate across extended development runs where Sonnet 5 might exit early.
Specialized domain capability marks another sharp divide between the models. Claude Sonnet 5 ships with deliberate cyber safeguards enabled by default and possesses significantly lower aptitude on cybersecurity tasks compared to Opus. While Sonnet 5 handles agentic tool execution—including operating terminals and browsers—at performance levels approaching Opus 4.8, it is intentionally restricted from offensive security workflows, vulnerability research, and penetration analysis. Claude Opus 5, conversely, delivers frontier-tier deep reasoning and specialized cybersecurity capabilities. It also offers dedicated integrations within consumer tools, serving Claude.ai, Claude Code, and Claude Cowork, whereas Sonnet 5 sits as the default daily driver for standard Free and Pro tier Claude accounts while supporting Team, Enterprise, and Max plans.
Best use case for Claude Sonnet 5
Claude Sonnet 5 suits developers and organizations running high-volume, cost-conscious applications that demand fast response times and solid agentic tool use without deep cybersecurity requirements.
Best use case for Claude Opus 5
Claude Opus 5 suits enterprise teams tackling complex software engineering, long-horizon autonomy, and cybersecurity challenges where maximum reasoning depth outweighs higher token costs and moderate base latency.
Decision framework
Choose Claude Sonnet 5 if your product demands high request throughput, rapid interactive latency, and cost-controlled token consumption. It is ideal for live conversational assistants, continuous data transformation, customer-facing chat, and agentic workflows that interact with browsers or command lines without specialized cybersecurity needs. At $2 in and $10 out per million tokens, Sonnet 5 lets you scale autonomous loops without compounding infrastructure bills.
Choose Claude Opus 5 if your system tackles complex, multi-turn software engineering tasks, cybersecurity research, and extended autonomous problem solving where failure carries high downstream costs. If your application requires deep architectural evaluation, automated code refactoring via Claude Code, or long-horizon tasks that require rigorous self-verification, the $5 in and $25 out token rates provide an intelligence tier designed to prevent the need for manual developer intervention.
Bottom line
Claude Sonnet 5 gives you a fast, cost-efficient workhorse priced at $2 in and $10 out per million tokens for everyday production throughput, whereas Claude Opus 5 gives you a frontier reasoning engine priced at $5 in and $25 out per million tokens built for complex autonomy and cybersecurity tasks. While both tools deploy identical 1M-token context buffers and cloud availability across the Claude API, AWS Bedrock, Google Cloud, and Microsoft Foundry, they should not be treated as interchangeable endpoints. Teams operating customer-facing interfaces, high-frequency tool pipelines, and latency-sensitive features will find Sonnet 5 far easier to sustain financially and operationally. Conversely, engineering departments deploying agents for multi-file refactoring, vulnerability inspection, and high-effort reasoning should absorb the cost of Opus 5, reserving Sonnet for the surrounding orchestration layers.
Sources and verification
The product facts have been checked against the sources below. The AI-assisted analysis was audited against these exact evidence records and approved by a human editor.
Last verified September 21, 2026
Last verified September 21, 2026
Editorial validation
Human-approvedApproved September 21, 2026 after an automated evidence audit using gemini-3.6-flash.
Read our comparison methodology and editorial policy, learn about TerraNet, or report a correction.
Common questions
No. Sending non-default sampling parameters such as temperature, top_p, or top_k to Claude Sonnet 5 will return a 400 error. The model uses adaptive thinking by default and requires developers to tune output behavior through Anthropic defined effort levels ranging from low to max.
Fast mode is a research preview available exclusively on the Claude API, running at approximately 2.5 times the default Opus 5 speed. It is not supported on partner clouds like AWS Bedrock or Google Cloud, nor is it available on the Batch API. Fast mode doubles the Opus 5 usage price to $10 per million input tokens and $50 per million output tokens.
Claude Sonnet 5 has cyber safeguards enabled by default and possesses significantly lower ability to execute cybersecurity tasks compared to the Opus tier. Teams working on vulnerability testing, exploit analysis, or specialized offensive security workflows require Claude Opus 5, which retains deep reasoning capabilities for that domain.
Yes. Both models provide a 1M-token input context window and up to 128K output tokens on the standard synchronous Messages API. Both models can also generate up to 300K output tokens when executing requests through the Message Batches API using the output-300k-2026-03-24 beta header.
AI-assisted draft audited against the cited product evidence and approved by a human editor. Vendor pricing and capabilities can change after the recorded verification date.
Continue researching
Claude Opus 5 pairs a 1M-token context window and 128K-token output ceiling with adaptive thinking enabled by default, serving as Anthropic's recommended starting model for agentic coding and deep reasoning at $5 per million input tokens and $25 per million output tokens. However, its Moderate comparative latency rating positions it behind faster options in latency-sensitive pipelines, and its breaking changes—which keep thinking enabled unless manually turned down at effort high or below—can disrupt production configurations carried over from older versions. Furthermore, its research-preview fast mode doubles the token rates and remains restricted to the Claude API rather than third-party cloud environments, while organizations with sovereign infrastructure requirements cannot self-host its closed weights. These technical constraints, pricing structures, and runtime realities lead engineering teams to explore alternatives across Anthropic's portfolio, hyperscaler competitors, and open-weight architectures.
Read guideClaude Sonnet 5 returns a 400 error whenever an API request supplies non-default sampling parameters like temperature, top_p, or top_k, or attempts to set a manual thinking token budget. This strict parameter enforcement invalidates existing prompt-engineering harnesses tuned for earlier generations and restricts runtime control strictly to an adaptive effort parameter. In addition, Anthropic deploys Sonnet 5 with cyber safeguards active by default, deliberately lowering its performance on cybersecurity tasks compared to Opus-tier models. Combined with a tokenizer that yields roughly thirty percent more tokens for identical text relative to Sonnet 4.6, technical teams often need alternative models that support legacy parameter overrides, offer specialized security capabilities, run on different provider clouds, or deliver significantly lower inference costs for high-volume pipelines.
Read guideClaude Haiku 4.5 provides low-latency execution and high-volume cost efficiency at $1 per million input tokens, while Claude Sonnet 5 provides a 1M-token context window and autonomous tool planning at double the base token price. Teams optimizing for interactive user experiences, live customer support desks, and narrow margin footprints will find Haiku 4.5 the more practical fit. Conversely, projects requiring broad document synthesis, deep programmatic refactoring, and independent multi-turn agent loops will find Sonnet 5 essential despite its higher token counts and strict 400-error validation on sampling overrides.
Read guideClaude Opus 5 costs half as much as Claude Fable 5.1 on standard input and output tokens while delivering faster response times and flexible reasoning toggles, whereas Claude Fable 5.1 delivers Anthropic's deepest reasoning capabilities alongside slower comparative latency and strict programmatic constraints. For standard agentic engineering and general enterprise workloads, Opus 5 provides the more balanced operational foundation due to its $5 and $25 token rates, optional 2.5-times fast mode, and ability to disable thinking when latency matters. Claude Fable 5.1 belongs in pipelines where evaluations demonstrate that Opus 5 cannot resolve the underlying reasoning problem, provided the engineering stack can accommodate double the token expense, mandatory adaptive thinking, and an API contract that prohibits forced tool selection.
Read guideGLM-5.3-Flash delivers an accessible, lightweight 18B active-parameter architecture under an unencumbered MIT license with rock-bottom operational costs, whereas Kimi K3 demands a massive multi-node 2.8T-parameter footprint, custom legal approval, and always-on reasoning tokens to unlock frontier-level agentic task completion. For high-volume production, multimodal file pipelines, and internal deployments on standard server setups, GLM-5.3-Flash provides the most practical and legally clear path forward. For demanding terminal control, automated web browsing, and multi-step programmatic problem solving where accuracy supersedes operational cost, Kimi K3 stands as the superior agentic reasoning tool.
Read guideGPT-5.6 Sol gives developers a high-throughput endpoint with switchable reasoning and low base token prices, whereas Claude Fable 5.1 provides an agentic reasoning engine with mandatory thinking and deeply discounted prompt caching across multi-cloud infrastructure. Organizations managing fast, high-volume production queues will find Sol's $4.00 entry rate and zero-effort option ideal for maintaining strict budget and latency targets. Conversely, enterprise teams managing complex analytical problems across recurring reference contexts will gain superior stability and sustainable long-term economics from Fable 5.1.
Read guide