Editorial illustration for Claude's four tiers explained: Fable, Opus, Sonnet and Haiku, and how to run them as a team
AI analysis / Latest briefings

Claude's four tiers explained: Fable, Opus, Sonnet and Haiku, and how to run them as a team

Anthropic now sells four Claude tiers at four prices, from $1 to $10 per million input tokens. This explainer says what each one is for, what the public catalogue says each scores, and draws four agent architectures in which the expensive model plans and the cheap ones do the volume.

By TerraNet Technologies23 min read28 sources
Editorial illustration for Claude's four tiers explained: Fable, Opus, Sonnet and Haiku, and how to run them as a team
claude models
fable vs opus vs sonnet vs haiku
anthropic model tiers
multi-agent systems
model routing
llm cost optimization

Watch

Claude's Four Tiers Explained: Fable, Opus, Sonnet and Haiku

Listen to this article

~23 min spoken. Keeps playing while you work in another tab.

Since 1 September 2026 there have been four current Claude models, one per tier: Claude Fable 5.1 at the top, Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5. They share one API, one tool-use format and, for the top three, one 1-million-token context window. What separates them is how much each one can do on its own, how long it will work unattended, and what a million tokens costs. The four are meant to be used together: the reason a company has a Haiku is so that an Opus does not have to read your logs.

This is a guide for people choosing between them or wiring them into one system. The prices and scores come from the same public catalogue our benchmark articles draw on - Artificial Analysis for prices and its index scores, Epoch AI for the rest - as they stood on 10 September 2026. The architectures are the ones Anthropic documents and, in two cases, has measured.

The ladder in one picture

Each step down the ladder gives up some score for a price that falls faster than the score does. That is the whole design, and it is why the tiers are worth combining rather than picking one.

Artificial Analysis Intelligence Index against price, across the Claude tiersArtificial Analysis Intelligence Index against blended price per million tokens: Claude Fable 5.1 53.4% at $20, Claude Opus 5 50.7% at $10, Claude Sonnet 5 38.4% at $4.00, Claude Haiku 4.5 17.6% at $2.00.Artificial Analysis Intelligence Index against price, acrossthe Claude tiersClaude Fable 5.1Claude Opus 5Claude Sonnet 5Claude Haiku 4.50%25%50%75%100%$2.00$5.00$10$20blended price per 1M tokens, log scaleClaude Fable 5.1 · 53.4% · $20/1M blendedClaude Fable 5.153.4% · $20 per 1MClaude Opus 5 · 50.7% · $10/1M blendedClaude Opus 550.7% · $10 per 1MClaude Sonnet 5 · 38.4% · $4.00/1M blendedClaude Sonnet 538.4% · $4.00 per 1MClaude Haiku 4.5 · 17.6% · $2.00/1M blendedClaude Haiku 4.517.6% · $2.00 per 1MSource: Artificial Analysis, https://artificialanalysis.ai/ Licence Free API, attribution required. Retrieved2026-09-10.Prices are list prices per million tokens from Artificial Analysis (https://artificialanalysis.ai/),attribution required.Blended price is 3:1 input to output. Each score is the model's best over the effort settings the source ran.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Model Artificial Analysis Intelligence Index Blended $/1M
Claude Fable 5.1 53.4% (max) $20
Claude Opus 5 50.7% (max) $10
Claude Sonnet 5 38.4% (max) $4.00
Claude Haiku 4.5 17.6% (reasoning) $2.00

The Artificial Analysis Intelligence Index is a composite across ten evaluations, so it flattens the differences a single benchmark would show; the capability ladder further down separates them. But the shape survives every benchmark we looked at: Fable 5.1 and Opus 5 are close, Sonnet 5 is a clear step behind, and Haiku 4.5 is in a different class - and also a different price.

What each tier is

Claude Fable 5.1: the model for work nothing else finishes

Fable is the tier Anthropic introduced in June 2026 above Opus, and Fable 5.1, released on 1 September, is the current model. Anthropic describes it as its most capable model for "features that span an entire codebase, code review, performance work, and multi-day autonomous sessions" - long-horizon runs, first-shot implementations of well-specified systems, and end-to-end deliverables such as financial analysis, spreadsheets and slide decks. Its API behaves differently from the rest of the line in ways that matter to a builder:

  • Thinking is always on. There is no setting to turn it off; you control depth with the effort parameter, from low to max.
  • Turns are long. A single request on a hard task can run for many minutes - Anthropic's migration guide calls a 15-minute request normal at high effort - so callers stream, set timeouts accordingly, and check in on runs rather than blocking.
  • It can decline. Safety classifiers for biology and most cybersecurity content can return a successful response whose stop reason is refusal. Anthropic provides a server-side fallback parameter that re-runs the request on another model in the same call; builders are expected to turn it on.
  • It delegates well. Anthropic's own guidance for Fable 5.1 is the opposite of what it was for earlier models: rather than suppressing sub-agents, use them freely and let them run asynchronously.

Fable 5.1 lists at $10 per million input tokens and $50 per million output tokens - double Opus - with cache reads at $0.25 per million, a quarter of Fable 5's rate, and batch requests at half price.

Claude Opus 5: the default for agents

Opus 5, released in July 2026, is the model Anthropic's own cost guidance says to start with for most agent workloads. It has the same 1M context, the same 128K output ceiling and the same effort ladder as Fable, at half the price ($5 in, $25 out). Thinking is on by default and can be turned off only at effort high or lower. It is the only tier with a fast mode, which runs the same model at up to two and a half times the output speed for Fable's price.

The catalogue puts Opus 5 within a few points of Fable 5.1 on most benchmarks, and Anthropic's own measurement is blunter: on a coding subset both models largely saturate, Opus 5 scored 91.7% to Fable 5's 91.3% at about 60% of the cost. The behavioural note that matters for a team is that Opus 5 reaches for sub-agents readily and verifies its own work without being asked, so prompts written for older models that beg it to delegate or to double-check now cause too much of both.

Claude Sonnet 5: the workhorse

Sonnet 5, released in June 2026, is the tier Anthropic positions as the best combination of speed and intelligence: near-Opus quality on coding and agentic work at $2 in and $10 out, with the 1M context window and the full effort ladder. Two details are worth knowing before adopting it. Its tokenizer produces roughly 30% more tokens than Sonnet 4.6 did for the same text, so budgets measured on the older model do not carry over. And it does not accept mid-conversation system messages, an operator channel the Fable and Opus tiers have.

Sonnet is where most production traffic should land by default: the tier you run a coding assistant, a support agent or a document pipeline on, and the tier you measure the others against.

Claude Haiku 4.5: the fast one

Haiku 4.5 is the oldest of the four, released in October 2025, and the cheapest by a wide margin at $1 in and $5 out. It is also the fastest in the catalogue's measurements, at about 101 output tokens per second against Sonnet 5's 83, Fable 5.1's 70 and Opus 5's 55. It has a 200K context window and a 64K output ceiling rather than 1M and 128K, and its thinking is configured the old way, with a token budget, because it predates the effort parameter.

Anthropic's measurement of where Haiku fits is specific: it answered knowledge questions at about a tenth of Opus 5's cost per question, at 63% accuracy against Opus 5's 92%. That is the profile of a model for high-volume work with checkable outputs - classification, extraction, routing, summarising what a bigger model will then read - and not for long agentic loops, where a cheap mistake early compounds into an expensive one late.

The four side by side

Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
API model id claude-fable-5-1 claude-opus-5 claude-sonnet-5 claude-haiku-4-5
Released 1 Sep 2026 24 Jul 2026 30 Jun 2026 15 Oct 2025
Context window 1M tokens 1M tokens 1M tokens 200K tokens
Max output 128K tokens 128K tokens 128K tokens 64K tokens
List price, per 1M tokens $10 in / $50 out $5 in / $25 out $2 in / $10 out $1 in / $5 out
Thinking Always on On by default; off only at effort high or below On by default; can be turned off Token budget (budget_tokens)
Effort levels low to max low to max low to max Not supported
Fast mode No Yes, at Fable's price No No
Anthropic's own framing Days-long autonomous work, the hardest problems first The default for agents Speed and intelligence for production traffic Fastest and cheapest, for simple tasks

What a million tokens costs

The price ladder is steeper than the capability ladder. Output tokens on Fable 5.1 cost ten times what they cost on Haiku 4.5, and each step is roughly a doubling, with Opus to Sonnet the biggest single drop.

What a million tokens costs on each Claude tierPer million tokens: Claude Fable 5.1 $10 in and $50 out, Claude Opus 5 $5.00 in and $25 out, Claude Sonnet 5 $2.00 in and $10 out, Claude Haiku 4.5 $1.00 in and $5.00 out.What a million tokens costs on each Claude tierinputoutput, US dollars per million tokens, log scale$1.00$2.00$5.00$10$20$50Claude Fable 5.1$20 blended 3:1Claude Fable 5.1 · input $10 per 1M tokensClaude Fable 5.1 · output $50 per 1M tokens$10$50Claude Opus 5$10 blended 3:1Claude Opus 5 · input $5.00 per 1M tokensClaude Opus 5 · output $25 per 1M tokens$5.00$25Claude Sonnet 5$4.00 blended 3:1Claude Sonnet 5 · input $2.00 per 1M tokensClaude Sonnet 5 · output $10 per 1M tokens$2.00$10Claude Haiku 4.5$2.00 blended 3:1Claude Haiku 4.5 · input $1.00 per 1M tokensClaude Haiku 4.5 · output $5.00 per 1M tokens$1.00$5.00Source: Artificial Analysis. Retrieved 2026-09-10.Prices are list prices per million tokens from Artificial Analysis (https://artificialanalysis.ai/),attribution required.Not shown: cache reads, batch requests and fast mode, which change the bill more than the list price does.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Model Input $/1M Output $/1M Blended 3:1 $/1M Output price vs Haiku 4.5
Claude Fable 5.1 $10 $50 $20 10×
Claude Opus 5 $5.00 $25 $10 5×
Claude Sonnet 5 $2.00 $10 $4.00 2×
Claude Haiku 4.5 $1.00 $5.00 $2.00 1×

List price is not what you pay. Three mechanisms change the bill more than the choice of tier, and all three work on every tier:

  • Prompt caching. A cached prefix - the system prompt, the tool definitions, a long document - is billed at a fraction of the input price on every request after the first, about a tenth on most tiers and less on Fable 5.1. For an agent that resends the same context every turn, this is the largest single saving available.
  • Batch processing. Anything nobody is waiting for can go through the Batch API at half price.
  • Effort. On the three tiers that support it, output_config.effort scales how much the model thinks and how many tool calls it makes. In Anthropic's runs, dropping Opus 5 from the default to medium gave up about 2 points on long-horizon coding for half the cost, and low gave up about 8 points for a quarter of it; on research and knowledge work the curve was nearly flat.

Anthropic's cost guide puts these in order: caching first, then input and output hygiene, then batch, then effort, and only then a change of model. That ordering is the frame for everything that follows.

What each tier scores

Six benchmarks on which all four tiers have an Artificial Analysis score, each bar the model's best result over the effort settings the source ran:

Four Claude tiers on 6 benchmarks6 benchmarks, four Claude tiers. On Artificial Analysis Intelligence Index: Fable 5.1 53.4%, Opus 5 50.7%, Sonnet 5 38.4%, Haiku 4.5 17.6%.Four Claude tiers on 6 benchmarksClaude Fable 5.1Claude Opus 5Claude Sonnet 5Claude Haiku 4.5Artificial Analysis Intelligence IndexFable 5.1Claude Fable 5.1 · Artificial Analysis Intelligence Index · 53.4% (max)53.4% (max)Opus 5Claude Opus 5 · Artificial Analysis Intelligence Index · 50.7% (max)50.7% (max)Sonnet 5Claude Sonnet 5 · Artificial Analysis Intelligence Index · 38.4% (max)38.4% (max)Haiku 4.5Claude Haiku 4.5 · Artificial Analysis Intelligence Index · 17.6% (reasoning)17.6% (reasoning)Artificial Analysis Coding IndexFable 5.1Claude Fable 5.1 · Artificial Analysis Coding Index · 81.6% (max)81.6% (max)Opus 5Claude Opus 5 · Artificial Analysis Coding Index · 78.0% (max)78.0% (max)Sonnet 5Claude Sonnet 5 · Artificial Analysis Coding Index · 71.5% (max)71.5% (max)Haiku 4.5Claude Haiku 4.5 · Artificial Analysis Coding Index · 43.9% (reasoning)43.9% (reasoning)GPQA Diamond (AA)Fable 5.1Claude Fable 5.1 · GPQA Diamond (AA) · 93.7% (max)93.7% (max)Opus 5Claude Opus 5 · GPQA Diamond (AA) · 93.7% (xhigh)93.7% (xhigh)Sonnet 5Claude Sonnet 5 · GPQA Diamond (AA) · 91.1% (max)91.1% (max)Haiku 4.5Claude Haiku 4.5 · GPQA Diamond (AA) · 67.2% (reasoning)67.2% (reasoning)Humanity's Last ExamFable 5.1Claude Fable 5.1 · Humanity's Last Exam · 59.1% (max)59.1% (max)Opus 5Claude Opus 5 · Humanity's Last Exam · 54.9% (max)54.9% (max)Sonnet 5Claude Sonnet 5 · Humanity's Last Exam · 41.3% (max)41.3% (max)Haiku 4.5Claude Haiku 4.5 · Humanity's Last Exam · 10.4% (reasoning)10.4% (reasoning)Terminal-Bench 2.1Fable 5.1Claude Fable 5.1 · Terminal-Bench 2.1 · 91.4% (max)91.4% (max)Opus 5Claude Opus 5 · Terminal-Bench 2.1 · 89.1% (max)89.1% (max)Sonnet 5Claude Sonnet 5 · Terminal-Bench 2.1 · 80.5% (max)80.5% (max)Haiku 4.5Claude Haiku 4.5 · Terminal-Bench 2.1 · 44.2% (reasoning)44.2% (reasoning)Long Context Reasoning (Artificial Analysis)Fable 5.1Claude Fable 5.1 · Long Context Reasoning (Artificial Analysis) · 85.3% (max)85.3% (max)Opus 5Claude Opus 5 · Long Context Reasoning (Artificial Analysis) · 82.0% (medium)82.0% (medium)Sonnet 5Claude Sonnet 5 · Long Context Reasoning (Artificial Analysis) · 82.0% (max)82.0% (max)Haiku 4.5Claude Haiku 4.5 · Long Context Reasoning (Artificial Analysis) · 74.3% (reasoning)74.3% (reasoning)Source: Artificial Analysis, https://artificialanalysis.ai/ Licence Free API, attribution required. Retrieved2026-09-10.Each bar is the model's best score over the effort settings the source ran; the setting is in brackets. †Relayed by the source from a vendor rather than run by it.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Benchmark Fable 5.1 Opus 5 Sonnet 5 Haiku 4.5
Artificial Analysis Intelligence Index 53.4% (max) 50.7% (max) 38.4% (max) 17.6% (reasoning)
Artificial Analysis Coding Index 81.6% (max) 78.0% (max) 71.5% (max) 43.9% (reasoning)
GPQA Diamond (AA) 93.7% (max) 93.7% (xhigh) 91.1% (max) 67.2% (reasoning)
Humanity's Last Exam 59.1% (max) 54.9% (max) 41.3% (max) 10.4% (reasoning)
Terminal-Bench 2.1 91.4% (max) 89.1% (max) 80.5% (max) 44.2% (reasoning)
Long Context Reasoning (Artificial Analysis) 85.3% (max) 82.0% (medium) 82.0% (max) 74.3% (reasoning)

Two things to notice. First, the gap between Fable 5.1 and Opus 5 is small everywhere: 53.4 against 50.7 on the Intelligence Index, 81.6 against 78.0 on the Coding Index, a tie at 93.7 on GPQA Diamond. On the evidence the catalogue holds, Fable's premium buys endurance and independence on long tasks more than it buys accuracy on short ones - which is exactly what Anthropic says it is for. Second, the gap between Sonnet 5 and Haiku 4.5 is not small anywhere: 71.5 against 43.9 on coding, 80.5 against 44.2 on Terminal-Bench 2.1. Haiku is in a different tier for a reason, and the architectures below treat it that way.

Our earlier piece on what leaderboards cannot tell you applies here in full: most of these sources publish no run-to-run interval, so a two-point gap is not a finding. The ranking of the four tiers is robust; the decimals are not.

Which tier for which job

The job Start with Why
Multi-day autonomous build, an unsolved engineering problem, an end-to-end deliverable Fable 5.1 Built for long-horizon work; Anthropic's advice is to give it your hardest problem first
An agent that codes, researches or operates tools for minutes at a time Opus 5 Anthropic's stated default for agents; near Fable on benchmarks at half the price
Production coding assistant, support agent, document pipeline Sonnet 5 Near-Opus quality at a fifth of Fable's price; the tier to measure the others against
Classification, extraction, routing, tagging, first-pass summaries at volume Haiku 4.5 Fastest and cheapest; strongest where the output can be checked
Planning or reviewing for a cheaper model that does the typing Fable 5.1 or Opus 5 as advisor or orchestrator The two multi-model shapes Anthropic has measured (below)
Anything nobody is waiting for The same tier, through the Batch API Half price on every tier

The honest caveat on this table is that a larger model at lower effort is often the cheaper option, and only a measurement on your own tasks can tell you. In Anthropic's runs, Fable 5 at low effort beat Sonnet 5 on a deep-research benchmark while costing about 10% less per task. Before you step down a tier, step down effort on the tier you have.

Using them together

An agent system built on one model pays that model's price for every token, including the tokens spent reading a log file or tagging a ticket. The point of a tiered line is that the work inside an agent is itself tiered: a little of it is hard judgement and most of it is reading, typing and checking. Anthropic's cost guide describes exactly two multi-model shapes it has measured to pay off - an advisor, where a cheap model runs the loop and consults an expensive one, and an orchestrator, where an expensive model plans and hands bulk work to cheap ones - and warns that both are architecture changes to be validated like one. The four diagrams below are those two shapes and two compositions of them.

1. Plan at the top, hand the work down

The pipeline most teams reach for first: the most capable model reads the codebase and writes the plan, a cheaper model implements it part by part, cheaper models still review and do the chores, and the planner checks the merged result against its own specification.

One feature, four tiers: plan at the top, hand the rest downA pipeline diagram: Claude Fable 5.1 plans and splits a feature, Claude Opus 5 implements each part, Claude Sonnet 5 reviews and writes tests while Claude Haiku 4.5 triages logs and fixtures, and Fable 5.1 verifies the merged result against its own spec, sending fix requests back to Opus 5.One feature, four tiers: plan at the top, hand the rest downClaude Fable 5.1Claude Opus 5Claude Sonnet 5Claude Haiku 4.5spec, per partdifflogs, fixturesfindings, testssummariesfix requestsCLAUDE FABLE 5.1Plan and decideReads thecodebase, writesthe spec, splitsit into parts,keeps the hardcalls.CLAUDE OPUS 5Implement eachpartWrites the codeand runs the testsfor one part ofthe spec at atime.CLAUDE SONNET 5Review and testAn independentpass over eachdiff; writes thetests the specasks for.CLAUDE HAIKU 4.5Bulk choresClassifies testfailures,summarises logs,extracts fixturesfrom many files.CLAUDE FABLE 5.1Verify againstthe specChecks the mergedchange againstwhat it asked forbefore it ships.Which model is the planner and which the implementer is a decision to measure, not a rule: on a coding subsetAnthropic measured, Opus 5 matched Fable 5 at about 60% of its cost, and lower effort on the same model is thefirst lever to try before changing tiers.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Step Model Does Hands off to
1. Plan Claude Fable 5.1 Reads the codebase, writes the spec, splits it into parts Opus 5, one part at a time
2. Implement Claude Opus 5 Writes and tests the code for each part Sonnet 5 (the diff), Haiku 4.5 (logs and fixtures)
3a. Review and test Claude Sonnet 5 Independent review of each diff; writes the tests the spec asks for Fable 5.1
3b. Bulk chores Claude Haiku 4.5 Classifies failures, summarises logs, extracts fixtures Fable 5.1
4. Verify Claude Fable 5.1 Checks the merged change against its own spec Opus 5 with fix requests, or ships

Why it is drawn this way:

  • Fable plans and verifies, Opus implements. Planning is where a wrong call is most expensive and where a model's ability to hold a whole codebase in view matters; implementation is a sequence of bounded tasks where Opus 5 is within a few points of Fable at half the cost. Verification returns to the planner because a fresh-context check against the spec catches what the implementer's self-review does not - Anthropic's Fable 5.1 guidance says separate verifier agents outperform self-critique.
  • Sonnet reviews, Haiku does chores. A review is judgement over a bounded diff, which is Sonnet's profile. Triaging 400 test failures into six causes, or pulling fixtures out of forty files, is volume with a checkable output, which is Haiku's.
  • The handoffs are files and briefs, not shared memory. Each model gets a self-contained brief with the paths, constraints and the report format it must return. Nothing downstream can see the planner's conversation.

In practice this pipeline is what a Claude Code session with sub-agents already is: the main session on one model, each sub-agent defined with its own model in its configuration, briefed and returning a report. The same shape runs on the Claude Agent SDK, and it is also what a Managed Agents roster produces when the lead is on Fable and the workers are on cheaper tiers.

2. A cheap executor with an expensive advisor

The inverse arrangement: the cheap model does all the work, and the expensive one is a consultant it can call on mid-turn.

The advisor pattern: Sonnet 5 runs the loop, Opus 5 is consultedA diagram of the advisor pattern: Claude Sonnet 5 runs every turn of the agent loop and every tool call, and consults a Claude Opus 5 advisor mid-turn only on hard decisions; the advice comes back into the same turn.The advisor pattern: Sonnet 5 runs the loop, Opus 5 isconsultedClaude Opus 5Claude Sonnet 5hard decisionadvicetool_usetool_resultRequestThe task, the system promptand your tool definitions.CLAUDE OPUS 5AdvisorConsulted mid-turn: theapproach, a stuck point, acheck before finishing.Billed only when asked.CLAUDE SONNET 5Executor runs the agentloopEvery ordinary turn andevery tool call: reads,edits, tests, searches.Your toolsYour code runs each toolcall and returns theresult.ResultMost of the tokens weregenerated at Sonnet prices.On the Claude API this is the advisor tool: the request names the executor as its model and the advisor insidethe tool definition. The advisor must be at least as capable as the executor, so the pairing runs up theladder, never down.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Step Model Does Hands off to
Executor Claude Sonnet 5 Runs the whole agent loop: every turn, every tool call The advisor, on a hard decision; your tools, on every tool call
Advisor Claude Opus 5 Answers a mid-turn question about approach or correctness The executor, which continues the same turn
Tools Your code Executes each tool call The executor

On the Claude API this is a single request with the advisor tool. The request's model is the executor; the advisor's model sits inside the tool definition, and the executor decides when to ask it:

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const response = await client.beta.messages.create({
  model: "claude-sonnet-5", // the executor: runs the loop and generates most of the tokens
  max_tokens: 16000,
  betas: ["advisor-tool-2026-03-01"],
  tools: [
    // Consulted only on hard decisions; capped so a hard task cannot run up the bill.
    { type: "advisor_20260301", name: "advisor", model: "claude-opus-5", max_uses: 3 },
    ...yourTools,
  ],
  messages: [{ role: "user", content: task }],
});

The rule the API enforces is that the advisor must be at least as capable as the executor: Sonnet 5 can consult Opus 5 or Fable 5.1, Opus 5 can consult Fable 5.1, but nothing can consult a model below it. Advice from an Opus 5 or Fable advisor comes back encrypted - the executor reads it, your code does not - which is a detail to know before you build a UI that wants to show it.

When it pays: when the capability gap is wide and the executor actually asks. Anthropic's measured warning is that the consult rate is fragile - lowering effort dropped one pairing from consulting on most tasks to almost none, at which point it scored below the executor alone - and that on its coding benchmark the flagship pairing was the most accurate configuration measured but sat within noise of the frontier model alone at medium effort, at about the same cost. Sweep effort and price the stronger model alone before adding an advisor.

3. An orchestrator with cheaper workers

The shape for work with bulk in it: many sources to read, many files to process, more material than one context window holds.

The orchestrator pattern: Opus 5 leads, Haiku 4.5 reads, Sonnet 5 reviewsA diagram of an orchestrator: a Claude Opus 5 lead briefs several Claude Haiku 4.5 researchers in parallel, a Claude Sonnet 5 reviewer, and a copy of itself for a large sub-analysis; only their reports return to the lead, which keeps verification and the synthesis for itself.The orchestrator pattern: Opus 5 leads, Haiku 4.5 reads,Sonnet 5 reviewsClaude Opus 5Claude Sonnet 5Claude Haiku 4.5one brief each, in parallelonly the reports returnCLAUDE OPUS 5LeadPlans the work, briefs eachworker with everything itneeds, verifies thereports, writes thesynthesis.CLAUDE HAIKU 4.5Researchers, several atonceOne well-scoped questioneach: search, read,extract, report withsources. Many input tokens,little hard reasoning.CLAUDE SONNET 5ReviewerAn independent pass over achange or a draft, withfile and line evidence.CLAUDE OPUS 5A copy of the leadOne large sub-analysis thatneeds the lead's fullcapability.ReportsEach worker runs in its owncontext window; only itsreport comes back, so thelead's context stays small.This is the shape a Managed Agents multiagent roster produces; the same shape can be built by hand with theClaude API by calling a cheaper model for each sub-task. It pays when there is bulk to hand off - manyindependent pieces, ideally more than one context window holds - and costs a plan, a handoff and a merge whenthere is not.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Step Model Does Hands off to
Lead Claude Opus 5 Plans, briefs, verifies, synthesises Researchers, the reviewer, a copy of itself
Researchers Claude Haiku 4.5 Search, read and extract for one question each, in parallel The lead, as a report with sources
Reviewer Claude Sonnet 5 Independent review with file and line evidence The lead
Copy of the lead Claude Opus 5 One sub-analysis needing full capability The lead

Managed Agents builds this from a roster. The lead is one stored agent; each worker is another, on its own model with a narrow system prompt and only the tools it needs; the lead's configuration lists them, and it decides when to spawn one. Anthropic's documented starting point is two agents:

worker = client.beta.agents.create(
    name="Web researcher",
    description="Fast, low-cost, read-only researcher. Give it one well-scoped question; it searches, reads, and reports findings with sources.",
    model="claude-haiku-4-5",
    system="Answer exactly the question you are given. Search and read as much as you need, then report concise findings with a source URL or file path for every claim.",
    tools=[{"type": "agent_toolset_20260401", "default_config": {"enabled": False},
            "configs": [{"name": n, "enabled": True} for n in ("read", "glob", "grep", "web_fetch", "web_search")]}],
)

lead = client.beta.agents.create(
    name="Research lead",
    description="Plans and synthesizes research.",
    model="claude-opus-5",
    system="Plan the work. Delegate each independent, reading-heavy question to Web researcher, one self-contained task per spawn, several in parallel. Keep verification and the final synthesis for yourself.",
    tools=[{"type": "agent_toolset_20260401"}],
    multiagent={"type": "coordinator", "agents": [worker.id, {"type": "self"}]},
)

Each worker runs in its own thread with a fresh context window, is billed at its own model's rates, and returns only its report, so the lead's context stays small however much the workers read. The constraints are worth knowing up front: one level of delegation, at most 20 roster entries and 25 concurrent threads, and threads share the container's files but not each other's conversation - so every brief has to carry the paths and the report format the worker needs.

When it pays: Anthropic measured the orchestrator at 55% less than the frontier model alone on work larger than any context window, three to seven points below the frontier model's best score. On routine search it paid as tail insurance - about half the average cost - and reversed on the harder full set. When the work is one dependent chain that fits in a single context, the orchestrator pays for a plan, a handoff and a merge that a single model gets for free, and in every such case Anthropic measured, the lead's model alone at lower effort came out ahead.

4. Route by difficulty, escalate on failure

The production shape for mixed traffic: a cheap first look decides which tier handles each request, and a checker that is not a model sends failures back up the ladder.

The router pattern: Haiku 4.5 triages, failures climb one tierA routing diagram: Claude Haiku 4.5 classifies each incoming request and routes routine work to Haiku 4.5, standard work to Claude Sonnet 5 and hard work to Claude Opus 5; a checker validates the output and sends failures back to the router to be escalated one tier.The router pattern: Haiku 4.5 triages, failures climb onetierClaude Opus 5Claude Sonnet 5Claude Haiku 4.5failed: escalate one tier, or raise effortIncoming requestA support ticket,a document, acoding task, aquery.CLAUDE HAIKU 4.5Classify androuteA cheap firstlook: which kindof task, how hard,what is missing.CLAUDE HAIKU 4.5RoutineExtraction,tagging, atemplated reply.Checkable output.CLAUDE SONNET 5StandardA bounded codechange, an answerthat needsjudgement.CLAUDE OPUS 5HardMulti-step agenticwork, an ambiguousspec, a longcontext.CheckerTests, a schema, arubric: somethingthat can say"wrong" without amodel.The checker is what makes this safe: without a signal that does not come from a model, the cheap tier has torecognise the cases it cannot handle, which is the judgement it is missing. Anthropic's measured version - runeverything cheap, re-run failures at the default - held the pass rate for about half the cost.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Step Model Does Hands off to
Router Claude Haiku 4.5 Classifies the request and picks a tier Haiku 4.5, Sonnet 5 or Opus 5
Routine Claude Haiku 4.5 Extraction, tagging, templated replies The checker
Standard Claude Sonnet 5 Bounded changes and answers needing judgement The checker
Hard Claude Opus 5 Multi-step agentic work and ambiguous specs The checker
Checker Your code Tests, schema validation or a rubric The router, on failure, to escalate

The checker is what makes the router safe. Without a signal that does not come from a model - tests, a JSON schema, a rubric with a threshold - the cheap tier has to recognise the cases it cannot handle, which is the very judgement it lacks. With one, the cascade can be measured. Anthropic's version of this is the simplest possible: run everything at low effort and re-run failures at the default. On its coding runs that passed about 93% of tasks for about $0.70 each, against 91.7% for $1.39 running everything at the default - the same pass rate for half the cost, counting the failed cheap attempts. The same policy works across tiers as it does across effort levels, with the same requirement: something has to be able to say "wrong".

Two implementation notes. Escalation should climb one step at a time - a Haiku failure goes to Sonnet, not to Fable - because the next tier up usually suffices and the price doubles at each step. And the routing decision itself is cheap only if it is small: a short classification with a structured output, not a conversation.

When one model beats the team

Everything above is conditional, and Anthropic's own guidance is unusually direct about the conditions. Two models beat one in exactly the two measured shapes, and both lose when their precondition is missing: an advisor without a wide capability gap or a consult it actually makes, an orchestrator without bulk to hand off. A single model at lower effort is the comparison every multi-model design has to beat, and it often does: on the newest models low frequently exceeds the xhigh or max result of the previous generation.

There is also a cost that never shows on a per-token price list. Prompt caches are per model, so a cascade forfeits the cache reuse a single model enjoys, and every handoff re-establishes context that the receiving model then has to read. Judge designs by cost per completed task, not per request - a cheaper request that needs two more turns to finish is not cheaper.

Practical rules

  1. Start on Opus 5 for agents and Sonnet 5 for everything else, at the default effort. Measure before you move.
  2. Turn on prompt caching and route anything asynchronous through the Batch API before touching the model. These are free wins on every tier.
  3. Sweep effort on the tier you have before stepping down a tier. Re-sweep after a model change; the curve is per workload and per model.
  4. Give Haiku 4.5 the volume, and only where the output can be checked: classification, extraction, routing, first-pass reading.
  5. Reserve Fable 5.1 for the work that does not finish on Opus: multi-day runs, first-shot builds of well-specified systems, the hardest unsolved problem you have. Stream, plan for long turns, and turn on the refusal fallback.
  6. When you combine tiers, pick one of the two measured shapes, write every handoff as a self-contained brief, and put a non-model checker at the end.
  7. Compare designs on cost per completed task on your own traffic, with repeated trials. Our model-selection pilot exists because a leaderboard cannot do that step for you.

Resources

Anthropic on the models

Cost, effort and thinking

Building the teams

The data behind the figures

AI Tools

    Claude's four tiers explained: Fable, Opus, Sonnet and Haiku, and how to run them as a team | TerraNet Technologies