All videos

Video briefing

Claude's Four Tiers Explained: Fable, Opus, Sonnet and Haiku

Claude Fable 5.1 · Claude Opus 5 · Claude Sonnet 5 · Claude Haiku 4.5 8:10

Anthropic offers four Claude tiers from $1 to $10 per million input tokens. Here is what public benchmarks show each model achieves and how to deploy them together.

Transcript

What the video says

Claude's four tiers explained: Fable, Opus, Sonnet and Haiku, and how to run them as a team. From TerraNet Technologies. If this helps, a like and a subscribe go a long way. Anthropic sells four Claude tiers at four different prices, spanning from $1 to $10 per million input tokens.

Rather than picking a single tier for every workload, the architecture pays off when you run them together, letting expensive models plan while cheap models handle volume across your systems. Each step down the ladder gives up some score for a price that falls faster than the score does.

Figure 1 plots the Artificial Analysis Intelligence Index against blended price per million tokens. Ranked from highest to lowest, Claude Fable 5.1 leads at 53.4 percent at $20. Claude Opus 5 follows at 50.7 percent at $10. Claude Sonnet 5 reaches 38.4 percent at $4.00, while Claude Haiku 4.5 marks 17.6 percent at $2.00. The horizontal axis measures blended cost per million tokens while the vertical axis tracks composite intelligence.

The takeaway is that benchmark score decreases gradually while price falls significantly faster at every downward step. Figure 2 tracks input and output token pricing across the Claude family, indexed against Haiku as a baseline. While input costs scale gradually, output pricing ramps aggressively. Sonnet doubles Haiku's baseline to ten dollars per million tokens, while Opus climbs to twenty-five dollars—a five-times jump. At the extreme, flagship Fable hits fifty dollars, commanding a ten-times premium on generated text.

The clear dynamic here is margin: vendors are heavily taxing generative generation over contextual processing, meaning agentic workloads with heavy output face a punishing cost penalty at higher tiers. Figure 3 maps capability across all four tiers, revealing a sharp non-linear drop between mainstream and entry models. At the frontier, Fable 5.1 and Opus 5 remain tightly matched—both clearing ninety-three percent on GPQA Diamond and leading Humanity’s Last Exam above fifty percent. Sonnet 5 holds solid ground in coding and general intelligence, trailing closely behind. The real cliff occurs at Haiku 4.5. Its performance collapses to ten percent on tough academic reasoning and drops by half on Terminal-Bench, proving efficiency at the bottom tier comes at the cost of deep reasoning.

List price does not dictate your final invoice. Prompt caching provides the largest saving on every tier by drastically discounting repeated prefixes. Asynchronous tasks routed through the Batch API run at half price. Adjusting the effort parameter allows teams to trim thinking cycles: Anthropic observed Opus 5 at low effort surrender about 8 points on coding tasks while reducing cost to a quarter, whereas research workloads showed a nearly flat curve. Anthropic's cost guide puts these steps in strict sequence: caching first, then input and output hygiene, then batch processing, then effort tuning, and only then a change of model.

Figure 4 outlines the pipeline architecture across four sequential stages. Claude Fable 5.1 starts at Step 1 to read code, write specifications, and split tasks. Claude Opus 5 receives each slice at Step 2 to write and test code. At Step 3, Claude Sonnet 5 conducts independent review while Claude Haiku 4.5 handles bulk chores like summarizing logs and extracting fixtures. At Step 4, Fable 5.1 verifies the merged deliverable against its spec, issuing rework back to Opus 5. The takeaway is that expensive models define and verify specifications while intermediate tiers execute bounded implementation and triage.

Figure 5 maps the advisor pattern across three operational roles. Claude Sonnet 5 acts as executor, running the primary agent loop, driving every turn, and dispatching tool calls. Claude Opus 5 serves as advisor, answering mid-turn inquiries regarding technical approach or correctness. Deterministic tools execute actions and return raw outputs directly back to Sonnet 5, which resumes its active turn without losing state. The chart organizes roles by execution layer and handoff destination. The primary takeaway is that an economical executor runs continuous agent loops while consulting a premium tier only for targeted mid-turn advice.

Figure 6 displays the orchestrator pattern across four specialized roles. Claude Opus 5 leads the architecture, planning objectives, coordinating briefs, and synthesizing final results. Multiple Claude Haiku 4.5 instances operate in parallel as researchers to search and extract sources. Claude Sonnet 5 acts as independent reviewer verifying evidence with citations, while an identical Opus 5 instance executes complex sub-analyses. All worker outputs return directly to the lead. The takeaway is that a high-tier orchestrator maintains context and synthesis while cheap parallel workers perform large-scale information retrieval and validation.

Figure 7 illustrates routing by difficulty and escalating on failure across four tiers of execution. Claude Haiku 4.5 acts as the router classifying each incoming request. It directs routine chores back to Claude Haiku 4.5, sends standard tasks needing judgement to Claude Sonnet 5, and assigns hard multi-step problems to Claude Opus 5. The diagram organizes stages from inbound classification to verification checks. A deterministic checker validates results against tests or schemas, routing failures back for escalation. The takeaway is that cheap models filter volume while escalating only when tasks fail validation.

Multi-model cascades carry trade-offs that per-token lists obscure. Because prompt caches do not transfer across different tiers, every handoff forces the receiving model to re-ingest context, incurring extra token charges. Furthermore, Anthropic demonstrated that a single advanced tier at low effort can outperform previous generation models running at maximum effort.

In Anthropic's runs, Fable 5 at low effort beat Sonnet 5 on a deep-research benchmark while costing about 10% less per task. Unless an orchestrator offloads massive volume or an advisor bridges a wide capability gap, a single tuned tier frequently finishes work cheaper.

Begin agent pipelines on Opus 5 and standard tasks on Sonnet 5 at default effort. Enable prompt caching and batch routes before modifying models, and sweep effort settings thoroughly on your selected tier. Haiku 4.5 should handle high volume only where outputs can be mechanically checked. Reserve Fable 5.1 for the multi-day runs and engineering challenges that fail to complete on Opus 5. Compare designs on cost per completed task on your own traffic with repeated trials before committing to complex architectures.

Review the comprehensive evaluation data, pricing formulas, and agent templates directly in our published technical report on Terranet Technologies.

Produced by TerraNet Technologies from the cited evidence behind the written article. Facts can change after the recorded date.