Claude's four tiers explained: Fable, Opus, Sonnet and Haiku, and how to run them as a team
Anthropic now sells four Claude tiers at four prices, from $1 to $10 per million input tokens. This explainer says what each one is for, what the public catalogue says each scores, and draws four agent architectures in which the expensive model plans and the cheap ones do the volume.
Watch
Claude's Four Tiers Explained: Fable, Opus, Sonnet and Haiku
~23 min spoken. Keeps playing while you work in another tab.
Since 1 September 2026 there have been four current Claude models, one per tier: Claude Fable 5.1 at the top, Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5. They share one API, one tool-use format and, for the top three, one 1-million-token context window. What separates them is how much each one can do on its own, how long it will work unattended, and what a million tokens costs. The four are meant to be used together: the reason a company has a Haiku is so that an Opus does not have to read your logs.
This is a guide for people choosing between them or wiring them into one system. The prices and scores come from the same public catalogue our benchmark articles draw on - Artificial Analysis for prices and its index scores, Epoch AI for the rest - as they stood on 10 September 2026. The architectures are the ones Anthropic documents and, in two cases, has measured.
The ladder in one picture
Each step down the ladder gives up some score for a price that falls faster than the score does. That is the whole design, and it is why the tiers are worth combining rather than picking one.
The numbers behind this chart
| Model | Artificial Analysis Intelligence Index | Blended $/1M |
|---|---|---|
| Claude Fable 5.1 | 53.4% (max) | $20 |
| Claude Opus 5 | 50.7% (max) | $10 |
| Claude Sonnet 5 | 38.4% (max) | $4.00 |
| Claude Haiku 4.5 | 17.6% (reasoning) | $2.00 |
The Artificial Analysis Intelligence Index is a composite across ten evaluations, so it flattens the differences a single benchmark would show; the capability ladder further down separates them. But the shape survives every benchmark we looked at: Fable 5.1 and Opus 5 are close, Sonnet 5 is a clear step behind, and Haiku 4.5 is in a different class - and also a different price.
What each tier is
Claude Fable 5.1: the model for work nothing else finishes
Fable is the tier Anthropic introduced in June 2026 above Opus, and Fable 5.1, released on 1 September, is the current model. Anthropic describes it as its most capable model for "features that span an entire codebase, code review, performance work, and multi-day autonomous sessions" - long-horizon runs, first-shot implementations of well-specified systems, and end-to-end deliverables such as financial analysis, spreadsheets and slide decks. Its API behaves differently from the rest of the line in ways that matter to a builder:
- Thinking is always on. There is no setting to turn it off; you control depth with the effort parameter, from
lowtomax. - Turns are long. A single request on a hard task can run for many minutes - Anthropic's migration guide calls a 15-minute request normal at high effort - so callers stream, set timeouts accordingly, and check in on runs rather than blocking.
- It can decline. Safety classifiers for biology and most cybersecurity content can return a successful response whose stop reason is
refusal. Anthropic provides a server-side fallback parameter that re-runs the request on another model in the same call; builders are expected to turn it on. - It delegates well. Anthropic's own guidance for Fable 5.1 is the opposite of what it was for earlier models: rather than suppressing sub-agents, use them freely and let them run asynchronously.
Fable 5.1 lists at $10 per million input tokens and $50 per million output tokens - double Opus - with cache reads at $0.25 per million, a quarter of Fable 5's rate, and batch requests at half price.
Claude Opus 5: the default for agents
Opus 5, released in July 2026, is the model Anthropic's own cost guidance says to start with for most agent workloads. It has the same 1M context, the same 128K output ceiling and the same effort ladder as Fable, at half the price ($5 in, $25 out). Thinking is on by default and can be turned off only at effort high or lower. It is the only tier with a fast mode, which runs the same model at up to two and a half times the output speed for Fable's price.
The catalogue puts Opus 5 within a few points of Fable 5.1 on most benchmarks, and Anthropic's own measurement is blunter: on a coding subset both models largely saturate, Opus 5 scored 91.7% to Fable 5's 91.3% at about 60% of the cost. The behavioural note that matters for a team is that Opus 5 reaches for sub-agents readily and verifies its own work without being asked, so prompts written for older models that beg it to delegate or to double-check now cause too much of both.
Claude Sonnet 5: the workhorse
Sonnet 5, released in June 2026, is the tier Anthropic positions as the best combination of speed and intelligence: near-Opus quality on coding and agentic work at $2 in and $10 out, with the 1M context window and the full effort ladder. Two details are worth knowing before adopting it. Its tokenizer produces roughly 30% more tokens than Sonnet 4.6 did for the same text, so budgets measured on the older model do not carry over. And it does not accept mid-conversation system messages, an operator channel the Fable and Opus tiers have.
Sonnet is where most production traffic should land by default: the tier you run a coding assistant, a support agent or a document pipeline on, and the tier you measure the others against.
Claude Haiku 4.5: the fast one
Haiku 4.5 is the oldest of the four, released in October 2025, and the cheapest by a wide margin at $1 in and $5 out. It is also the fastest in the catalogue's measurements, at about 101 output tokens per second against Sonnet 5's 83, Fable 5.1's 70 and Opus 5's 55. It has a 200K context window and a 64K output ceiling rather than 1M and 128K, and its thinking is configured the old way, with a token budget, because it predates the effort parameter.
Anthropic's measurement of where Haiku fits is specific: it answered knowledge questions at about a tenth of Opus 5's cost per question, at 63% accuracy against Opus 5's 92%. That is the profile of a model for high-volume work with checkable outputs - classification, extraction, routing, summarising what a bigger model will then read - and not for long agentic loops, where a cheap mistake early compounds into an expensive one late.
The four side by side
| Claude Fable 5.1 | Claude Opus 5 | Claude Sonnet 5 | Claude Haiku 4.5 | |
|---|---|---|---|---|
| API model id | claude-fable-5-1 |
claude-opus-5 |
claude-sonnet-5 |
claude-haiku-4-5 |
| Released | 1 Sep 2026 | 24 Jul 2026 | 30 Jun 2026 | 15 Oct 2025 |
| Context window | 1M tokens | 1M tokens | 1M tokens | 200K tokens |
| Max output | 128K tokens | 128K tokens | 128K tokens | 64K tokens |
| List price, per 1M tokens | $10 in / $50 out | $5 in / $25 out | $2 in / $10 out | $1 in / $5 out |
| Thinking | Always on | On by default; off only at effort high or below | On by default; can be turned off | Token budget (budget_tokens) |
| Effort levels | low to max | low to max | low to max | Not supported |
| Fast mode | No | Yes, at Fable's price | No | No |
| Anthropic's own framing | Days-long autonomous work, the hardest problems first | The default for agents | Speed and intelligence for production traffic | Fastest and cheapest, for simple tasks |
What a million tokens costs
The price ladder is steeper than the capability ladder. Output tokens on Fable 5.1 cost ten times what they cost on Haiku 4.5, and each step is roughly a doubling, with Opus to Sonnet the biggest single drop.
The numbers behind this chart
| Model | Input $/1M | Output $/1M | Blended 3:1 $/1M | Output price vs Haiku 4.5 |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $20 | 10× |
| Claude Opus 5 | $5.00 | $25 | $10 | 5× |
| Claude Sonnet 5 | $2.00 | $10 | $4.00 | 2× |
| Claude Haiku 4.5 | $1.00 | $5.00 | $2.00 | 1× |
List price is not what you pay. Three mechanisms change the bill more than the choice of tier, and all three work on every tier:
- Prompt caching. A cached prefix - the system prompt, the tool definitions, a long document - is billed at a fraction of the input price on every request after the first, about a tenth on most tiers and less on Fable 5.1. For an agent that resends the same context every turn, this is the largest single saving available.
- Batch processing. Anything nobody is waiting for can go through the Batch API at half price.
- Effort. On the three tiers that support it,
output_config.effortscales how much the model thinks and how many tool calls it makes. In Anthropic's runs, dropping Opus 5 from the default tomediumgave up about 2 points on long-horizon coding for half the cost, andlowgave up about 8 points for a quarter of it; on research and knowledge work the curve was nearly flat.
Anthropic's cost guide puts these in order: caching first, then input and output hygiene, then batch, then effort, and only then a change of model. That ordering is the frame for everything that follows.
What each tier scores
Six benchmarks on which all four tiers have an Artificial Analysis score, each bar the model's best result over the effort settings the source ran:
The numbers behind this chart
| Benchmark | Fable 5.1 | Opus 5 | Sonnet 5 | Haiku 4.5 |
|---|---|---|---|---|
| Artificial Analysis Intelligence Index | 53.4% (max) | 50.7% (max) | 38.4% (max) | 17.6% (reasoning) |
| Artificial Analysis Coding Index | 81.6% (max) | 78.0% (max) | 71.5% (max) | 43.9% (reasoning) |
| GPQA Diamond (AA) | 93.7% (max) | 93.7% (xhigh) | 91.1% (max) | 67.2% (reasoning) |
| Humanity's Last Exam | 59.1% (max) | 54.9% (max) | 41.3% (max) | 10.4% (reasoning) |
| Terminal-Bench 2.1 | 91.4% (max) | 89.1% (max) | 80.5% (max) | 44.2% (reasoning) |
| Long Context Reasoning (Artificial Analysis) | 85.3% (max) | 82.0% (medium) | 82.0% (max) | 74.3% (reasoning) |
Two things to notice. First, the gap between Fable 5.1 and Opus 5 is small everywhere: 53.4 against 50.7 on the Intelligence Index, 81.6 against 78.0 on the Coding Index, a tie at 93.7 on GPQA Diamond. On the evidence the catalogue holds, Fable's premium buys endurance and independence on long tasks more than it buys accuracy on short ones - which is exactly what Anthropic says it is for. Second, the gap between Sonnet 5 and Haiku 4.5 is not small anywhere: 71.5 against 43.9 on coding, 80.5 against 44.2 on Terminal-Bench 2.1. Haiku is in a different tier for a reason, and the architectures below treat it that way.
Our earlier piece on what leaderboards cannot tell you applies here in full: most of these sources publish no run-to-run interval, so a two-point gap is not a finding. The ranking of the four tiers is robust; the decimals are not.
Which tier for which job
| The job | Start with | Why |
|---|---|---|
| Multi-day autonomous build, an unsolved engineering problem, an end-to-end deliverable | Fable 5.1 | Built for long-horizon work; Anthropic's advice is to give it your hardest problem first |
| An agent that codes, researches or operates tools for minutes at a time | Opus 5 | Anthropic's stated default for agents; near Fable on benchmarks at half the price |
| Production coding assistant, support agent, document pipeline | Sonnet 5 | Near-Opus quality at a fifth of Fable's price; the tier to measure the others against |
| Classification, extraction, routing, tagging, first-pass summaries at volume | Haiku 4.5 | Fastest and cheapest; strongest where the output can be checked |
| Planning or reviewing for a cheaper model that does the typing | Fable 5.1 or Opus 5 as advisor or orchestrator | The two multi-model shapes Anthropic has measured (below) |
| Anything nobody is waiting for | The same tier, through the Batch API | Half price on every tier |
The honest caveat on this table is that a larger model at lower effort is often the cheaper option, and only a measurement on your own tasks can tell you. In Anthropic's runs, Fable 5 at low effort beat Sonnet 5 on a deep-research benchmark while costing about 10% less per task. Before you step down a tier, step down effort on the tier you have.
Using them together
An agent system built on one model pays that model's price for every token, including the tokens spent reading a log file or tagging a ticket. The point of a tiered line is that the work inside an agent is itself tiered: a little of it is hard judgement and most of it is reading, typing and checking. Anthropic's cost guide describes exactly two multi-model shapes it has measured to pay off - an advisor, where a cheap model runs the loop and consults an expensive one, and an orchestrator, where an expensive model plans and hands bulk work to cheap ones - and warns that both are architecture changes to be validated like one. The four diagrams below are those two shapes and two compositions of them.
1. Plan at the top, hand the work down
The pipeline most teams reach for first: the most capable model reads the codebase and writes the plan, a cheaper model implements it part by part, cheaper models still review and do the chores, and the planner checks the merged result against its own specification.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| 1. Plan | Claude Fable 5.1 | Reads the codebase, writes the spec, splits it into parts | Opus 5, one part at a time |
| 2. Implement | Claude Opus 5 | Writes and tests the code for each part | Sonnet 5 (the diff), Haiku 4.5 (logs and fixtures) |
| 3a. Review and test | Claude Sonnet 5 | Independent review of each diff; writes the tests the spec asks for | Fable 5.1 |
| 3b. Bulk chores | Claude Haiku 4.5 | Classifies failures, summarises logs, extracts fixtures | Fable 5.1 |
| 4. Verify | Claude Fable 5.1 | Checks the merged change against its own spec | Opus 5 with fix requests, or ships |
Why it is drawn this way:
- Fable plans and verifies, Opus implements. Planning is where a wrong call is most expensive and where a model's ability to hold a whole codebase in view matters; implementation is a sequence of bounded tasks where Opus 5 is within a few points of Fable at half the cost. Verification returns to the planner because a fresh-context check against the spec catches what the implementer's self-review does not - Anthropic's Fable 5.1 guidance says separate verifier agents outperform self-critique.
- Sonnet reviews, Haiku does chores. A review is judgement over a bounded diff, which is Sonnet's profile. Triaging 400 test failures into six causes, or pulling fixtures out of forty files, is volume with a checkable output, which is Haiku's.
- The handoffs are files and briefs, not shared memory. Each model gets a self-contained brief with the paths, constraints and the report format it must return. Nothing downstream can see the planner's conversation.
In practice this pipeline is what a Claude Code session with sub-agents already is: the main session on one model, each sub-agent defined with its own model in its configuration, briefed and returning a report. The same shape runs on the Claude Agent SDK, and it is also what a Managed Agents roster produces when the lead is on Fable and the workers are on cheaper tiers.
2. A cheap executor with an expensive advisor
The inverse arrangement: the cheap model does all the work, and the expensive one is a consultant it can call on mid-turn.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| Executor | Claude Sonnet 5 | Runs the whole agent loop: every turn, every tool call | The advisor, on a hard decision; your tools, on every tool call |
| Advisor | Claude Opus 5 | Answers a mid-turn question about approach or correctness | The executor, which continues the same turn |
| Tools | Your code | Executes each tool call | The executor |
On the Claude API this is a single request with the advisor tool. The request's model is the executor; the advisor's model sits inside the tool definition, and the executor decides when to ask it:
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.beta.messages.create({
model: "claude-sonnet-5", // the executor: runs the loop and generates most of the tokens
max_tokens: 16000,
betas: ["advisor-tool-2026-03-01"],
tools: [
// Consulted only on hard decisions; capped so a hard task cannot run up the bill.
{ type: "advisor_20260301", name: "advisor", model: "claude-opus-5", max_uses: 3 },
...yourTools,
],
messages: [{ role: "user", content: task }],
});
The rule the API enforces is that the advisor must be at least as capable as the executor: Sonnet 5 can consult Opus 5 or Fable 5.1, Opus 5 can consult Fable 5.1, but nothing can consult a model below it. Advice from an Opus 5 or Fable advisor comes back encrypted - the executor reads it, your code does not - which is a detail to know before you build a UI that wants to show it.
When it pays: when the capability gap is wide and the executor actually asks. Anthropic's measured warning is that the consult rate is fragile - lowering effort dropped one pairing from consulting on most tasks to almost none, at which point it scored below the executor alone - and that on its coding benchmark the flagship pairing was the most accurate configuration measured but sat within noise of the frontier model alone at medium effort, at about the same cost. Sweep effort and price the stronger model alone before adding an advisor.
3. An orchestrator with cheaper workers
The shape for work with bulk in it: many sources to read, many files to process, more material than one context window holds.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| Lead | Claude Opus 5 | Plans, briefs, verifies, synthesises | Researchers, the reviewer, a copy of itself |
| Researchers | Claude Haiku 4.5 | Search, read and extract for one question each, in parallel | The lead, as a report with sources |
| Reviewer | Claude Sonnet 5 | Independent review with file and line evidence | The lead |
| Copy of the lead | Claude Opus 5 | One sub-analysis needing full capability | The lead |
Managed Agents builds this from a roster. The lead is one stored agent; each worker is another, on its own model with a narrow system prompt and only the tools it needs; the lead's configuration lists them, and it decides when to spawn one. Anthropic's documented starting point is two agents:
worker = client.beta.agents.create(
name="Web researcher",
description="Fast, low-cost, read-only researcher. Give it one well-scoped question; it searches, reads, and reports findings with sources.",
model="claude-haiku-4-5",
system="Answer exactly the question you are given. Search and read as much as you need, then report concise findings with a source URL or file path for every claim.",
tools=[{"type": "agent_toolset_20260401", "default_config": {"enabled": False},
"configs": [{"name": n, "enabled": True} for n in ("read", "glob", "grep", "web_fetch", "web_search")]}],
)
lead = client.beta.agents.create(
name="Research lead",
description="Plans and synthesizes research.",
model="claude-opus-5",
system="Plan the work. Delegate each independent, reading-heavy question to Web researcher, one self-contained task per spawn, several in parallel. Keep verification and the final synthesis for yourself.",
tools=[{"type": "agent_toolset_20260401"}],
multiagent={"type": "coordinator", "agents": [worker.id, {"type": "self"}]},
)
Each worker runs in its own thread with a fresh context window, is billed at its own model's rates, and returns only its report, so the lead's context stays small however much the workers read. The constraints are worth knowing up front: one level of delegation, at most 20 roster entries and 25 concurrent threads, and threads share the container's files but not each other's conversation - so every brief has to carry the paths and the report format the worker needs.
When it pays: Anthropic measured the orchestrator at 55% less than the frontier model alone on work larger than any context window, three to seven points below the frontier model's best score. On routine search it paid as tail insurance - about half the average cost - and reversed on the harder full set. When the work is one dependent chain that fits in a single context, the orchestrator pays for a plan, a handoff and a merge that a single model gets for free, and in every such case Anthropic measured, the lead's model alone at lower effort came out ahead.
4. Route by difficulty, escalate on failure
The production shape for mixed traffic: a cheap first look decides which tier handles each request, and a checker that is not a model sends failures back up the ladder.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| Router | Claude Haiku 4.5 | Classifies the request and picks a tier | Haiku 4.5, Sonnet 5 or Opus 5 |
| Routine | Claude Haiku 4.5 | Extraction, tagging, templated replies | The checker |
| Standard | Claude Sonnet 5 | Bounded changes and answers needing judgement | The checker |
| Hard | Claude Opus 5 | Multi-step agentic work and ambiguous specs | The checker |
| Checker | Your code | Tests, schema validation or a rubric | The router, on failure, to escalate |
The checker is what makes the router safe. Without a signal that does not come from a model - tests, a JSON schema, a rubric with a threshold - the cheap tier has to recognise the cases it cannot handle, which is the very judgement it lacks. With one, the cascade can be measured. Anthropic's version of this is the simplest possible: run everything at low effort and re-run failures at the default. On its coding runs that passed about 93% of tasks for about $0.70 each, against 91.7% for $1.39 running everything at the default - the same pass rate for half the cost, counting the failed cheap attempts. The same policy works across tiers as it does across effort levels, with the same requirement: something has to be able to say "wrong".
Two implementation notes. Escalation should climb one step at a time - a Haiku failure goes to Sonnet, not to Fable - because the next tier up usually suffices and the price doubles at each step. And the routing decision itself is cheap only if it is small: a short classification with a structured output, not a conversation.
When one model beats the team
Everything above is conditional, and Anthropic's own guidance is unusually direct about the conditions. Two models beat one in exactly the two measured shapes, and both lose when their precondition is missing: an advisor without a wide capability gap or a consult it actually makes, an orchestrator without bulk to hand off. A single model at lower effort is the comparison every multi-model design has to beat, and it often does: on the newest models low frequently exceeds the xhigh or max result of the previous generation.
There is also a cost that never shows on a per-token price list. Prompt caches are per model, so a cascade forfeits the cache reuse a single model enjoys, and every handoff re-establishes context that the receiving model then has to read. Judge designs by cost per completed task, not per request - a cheaper request that needs two more turns to finish is not cheaper.
Practical rules
- Start on Opus 5 for agents and Sonnet 5 for everything else, at the default effort. Measure before you move.
- Turn on prompt caching and route anything asynchronous through the Batch API before touching the model. These are free wins on every tier.
- Sweep effort on the tier you have before stepping down a tier. Re-sweep after a model change; the curve is per workload and per model.
- Give Haiku 4.5 the volume, and only where the output can be checked: classification, extraction, routing, first-pass reading.
- Reserve Fable 5.1 for the work that does not finish on Opus: multi-day runs, first-shot builds of well-specified systems, the hardest unsolved problem you have. Stream, plan for long turns, and turn on the refusal fallback.
- When you combine tiers, pick one of the two measured shapes, write every handoff as a self-contained brief, and put a non-model checker at the end.
- Compare designs on cost per completed task on your own traffic, with repeated trials. Our model-selection pilot exists because a leaderboard cannot do that step for you.
Resources
Anthropic on the models
- Claude Fable 5.1 and Claude Mythos 5.1 - the announcement, and the system card.
- Introducing Claude Fable 5 and Claude Mythos 5 - where the tier above Opus came from.
- Introducing Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5.
- Models overview - model ids, context windows and output limits, and the Models API that returns them live.
- Choosing a model and the migration guide, which lists what breaks when you move between tiers.
Cost, effort and thinking
- Pricing - list prices, cache and batch rates, fast mode.
- Optimizing for cost and intelligence - the measured results quoted throughout this article, and the cookbook that runs them end to end.
- Effort, adaptive thinking, prompt caching, batch processing and context windows.
Building the teams
- Tool use overview, the advisor tool and programmatic tool calling.
- Managed Agents and its multi-agent orchestration guide, which the orchestrator example follows.
- Claude Agent SDK and Claude Code subagents, for the pipeline shape on your own machine.
- Building effective agents and How we built our multi-agent research system, Anthropic's engineering write-ups on when to build an agent at all and what an orchestrator costs.
The data behind the figures
- Artificial Analysis - prices, speeds and the index scores.
- Epoch AI's Benchmarking Hub - the independently run evaluations in the wider catalogue.
- Our own leaderboards and the noise floor, on how much of a benchmark gap is real.