claude opus vs sonnetclaude haiku vs sonnetwhich claude model to useclaude models

Claude Opus vs Sonnet vs Haiku: Which Tier to Use, and the Fourth One Nobody Compares

Anthropic answers the pick-a-model question three different ways on three of its own pages. Here is what the prices, the benchmarks and the plan gates actually say, including why the cheaper tier finished a whole benchmark for only 13% less.

Srdjan Bogicevic·
Claude Opus vs Sonnet vs Haiku: Which Tier to Use, and the Fourth One Nobody Compares

Claude's model picker has four names in it. Three of them are the ones everyone compares, and the fourth is the one Anthropic puts at the top of its own lineup page.

That is the first thing worth fixing about this question. Opus, Sonnet and Haiku describe a ladder that has since grown a rung, and Anthropic's own beginner tutorial is called "Choosing the right Claude model: Haiku, Sonnet, Opus, or Fable."

The second thing is stranger. Anthropic answers the pick-a-model question in three different places, and the three answers point in three different directions. Its models overview says to start at the top: "If you're unsure which model to use, start with Claude Opus 5 for most workloads." Its consumer tutorial says to start in the middle, calling Sonnet "your versatile default" and adding, "If you're not sure which model to pick, start here." And its guide on choosing a model opens with an efficiency-first option that begins at the bottom, with Haiku, and upgrades only where the cheaper model falls short.

All three are defensible. They answer different questions, and which one applies to you depends on what you are actually spending. I read all of them, plus the pricing tables, the deprecation schedule and Anthropic's published cost measurements, on 8 September 2026, and cross-checked the benchmark scores against Artificial Analysis the same day.

The number that surprised me most: running a full benchmark suite on Sonnet cost within 13% of running it on Opus, even though Sonnet's per-token price is 60% lower.

The Short Answer: Which Claude Model to Use

Tier Anthropic's own description API price per million tokens On claude.ai Reach for it when
Haiku "The fastest model with near-frontier intelligence" $1 in / $5 out Free and up The answer is short and you can check it at a glance
Sonnet "The best combination of speed and intelligence" $2 in / $10 out Free and up Writing, analysis, research, everyday work
Opus "For complex agentic coding and enterprise work" $5 in / $25 out Pro and up Reasoning you intend to push back on
Fable "For demanding reasoning and long-horizon agentic work" $10 in / $50 out Usage credits on Pro, inside the allowance on Max Long jobs you would rather describe than supervise

Before any of the detail, three things about that table.

The price ladder doubles at every step. One dollar, two, five, ten per million input tokens, and the same shape on output. Haiku to Fable is a tenfold gap on paper. It is not a tenfold gap in practice, in either direction, and most of this post is about why.

The tiers are not the same generation. Sonnet, Opus and Fable are current-generation models. Haiku is a version behind, and that shows up across four rows of the spec sheet, not one.

One tier is available without being included. Anthropic's pricing table marks Fable as "Usage credits" on Pro, so a Pro subscriber can select it but pays API rates for every message instead of drawing on the plan. For that one tier, "available" and "included" are different words.

The Three-Tier Question Now Has Four Answers

Fable is not a preview or an experiment. Anthropic's choosing-a-model guide calls it "Anthropic's most capable widely released model", it has its own row on the consumer pricing table, and it sits first in the comparison table on the models overview. Below it in the footer of claude.com sits a fifth name, Mythos, which the documentation describes as offering the same capabilities to a limited set of participants.

So the lineup is four models for anyone reading this, and the honest version of the question is Haiku vs Sonnet vs Opus vs Fable.

This matters more than a naming quibble, because the tier everyone treats as the ceiling is now the second rung from the top. Advice built around "Opus for the hard stuff" was written when Opus was the hard stuff. Anthropic's current escalation ladder reads differently. Its tutorial says to use Opus for "problems where you've tested with Sonnet and it struggled", and to use Fable for "problems you've tested with Opus and it struggled". Each tier is defined by the failure of the one below it, which is a more useful rule than any feature list.

The Spec Sheet, and the Rows That Actually Differ

Marketing copy for these tiers is mostly adjectives. The comparison table on Anthropic's models overview is not, and it is where the differences stop being vague.

Haiku Sonnet Opus Fable
Context window 200K tokens 1M tokens 1M tokens 1M tokens
Max output 64K tokens 128K tokens 128K tokens 128K tokens
Thinking Extended Adaptive Adaptive Adaptive, always on
Effort setting Not supported Yes, default high Yes, default high Yes, default high
Reliable knowledge cutoff Feb 2025 Jan 2026 May 2026 Jun 2026
Earliest retirement date 15 Oct 2026 30 Jun 2027 24 Jul 2027 1 Sep 2027
Latency, Anthropic's rating Fastest Fast Moderate Slower

Read down the Haiku column and the picture changes. Haiku is not simply a smaller Sonnet at half the price. It has a fifth of the context window, half the maximum output, no effort control, and a reliable knowledge cutoff sixteen months older than Fable's. Anthropic's tutorial puts it plainly enough, saying Haiku "rivals the reasoning capabilities" of a Sonnet release from the previous generation.

Two rows in that table are worth pulling out.

The retirement dates. Anthropic publishes a tentative earliest retirement date for every active model. Haiku's is 15 October 2026, which is about five weeks from the day I am writing this. Every other model in the current lineup has a date in 2027. To be fair to Anthropic, this is a floor rather than a schedule: Haiku is listed as Active with no deprecation notice, and the policy promises at least 60 days' warning before any public model retires. But if you are building a habit or a workflow around one tier, it is the one tier whose commitment expires this year.

The tokenizer. This one is buried in a footnote on the pricing page and it quietly rewrites the price comparison. Anthropic's newer models use a different tokenizer that "produces approximately 30% more tokens for the same text". Sonnet, Opus and Fable use it. Haiku does not. So the same paragraph of your writing bills as roughly 30% more tokens on the three current-generation models than it does on Haiku. The sticker gap between Haiku and Sonnet is 2x. Measured in words rather than tokens, it is closer to 2.6x, and the gap up to Fable is nearer 13x than 10x.

While you are looking at that table, one number on it does not survive contact with the chat app. Anthropic's consumer pricing table lists the context window as 200k on Free, Pro and both Max plans, for every model, which contradicts the 1M figure in the developer table and the help centre's own article on paid-plan context windows. I went through that contradiction in detail in the Claude limits post and will not relitigate it here. The practical version: treat the 1M window as an API property, not a promise about what your subscription gives you.

The Cost per Answer Does Not Follow the Price List

Here is the part that changed how I think about this question.

Artificial Analysis runs every major model through a fixed set of ten evaluations and publishes both the score and what the run cost. Because the task set is identical for every model, the total bill is a fair measure of what finishing the same work costs on each tier. These are the four Claude models as of 8 September 2026, on version 4.3 of that index.

Tier Intelligence index Rank Output speed Price per MTok Cost to finish the index Output tokens written
Fable 53 1st of 202 69/sec $10 / $50 $7.63 190M
Opus 51 7th of 202 54/sec $5 / $25 $5.86 140M
Sonnet 38 43rd of 202 80/sec $2 / $10 $5.09 370M
Haiku 18 139th of 202 83/sec $1 / $5 $0.21 78M

Sonnet's per-token price is 60% below Opus's. Its cost to finish the same benchmark is 13% below. The premium you thought you were avoiding mostly did not exist.

The explanation is in the last column. Sonnet wrote 370 million tokens to get through the index. Opus wrote 140 million. Sonnet is the most verbose model in the group by a wide margin, more than four times the 90-million median across all 202 models on the list, and every one of those extra tokens bills at the output rate. Cheaper per token, more tokens, similar bill.

Anthropic says the same thing in its own words, and it is the single most useful sentence in the whole cost guide: "Price lists are written per token, and per token the frontier model looks expensive. You pay for completed tasks, though, so compare models on cost per completed task." Its measured example is blunter than mine. On a long-horizon coding benchmark, Opus at its low effort setting solved 84% of tasks at $0.25 per solved task, while Sonnet at its default setting solved 77.4% at $0.84. The more expensive tier was more accurate and cheaper per result at the same time.

Two honest caveats before anyone rewires their workflow around this.

It does not hold everywhere. On a deep-research benchmark, the same comparison reverses hard: the frontier model scored ten points higher than Sonnet at roughly four times the cost per task, because on long research loops the more capable model does more work rather than less. Anthropic's conclusion is that "the ranking flips by workload, and no price list tells you which way."

And a benchmark suite is not your inbox. The index measures hard, structured problems. If your day is short questions with short answers, Sonnet's verbosity costs you very little, and the price list is a fine guide.

Claude Opus vs Sonnet: A Smaller Gap and a Smaller Saving Than You Think

On the numbers above, Opus scores 51 against Sonnet's 38, which is a real gap: thirteen points and thirty-six places on a 202-model list. It is not the gap between a serious model and a toy. Sonnet is still in the top quarter of everything measured.

What has changed is the price side of the trade. Sonnet used to cost $3 per million input tokens and $15 per million output. It now costs $2 and $10, and there is a small piece of news buried in that. Anthropic announced the lower price as introductory, due to expire on 31 August 2026, then published a line on the pricing page confirming that "the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur". The cut is permanent. Any comparison written before September that warns you about an upcoming Sonnet price rise is describing something that did not happen.

So Sonnet got cheaper per token, and still finishes hard work for only slightly less than Opus. Where does that leave the choice?

Use Sonnet when the work is bounded and you can tell at a glance whether the answer is right. Drafting, editing, summarising, straightforward analysis, question-and-answer over a document you know. Anthropic's tutorial lists writing, content creation, analysis that needs reasoning but is not extremely complex, and multi-step workflows, and that list matches my experience of where the extra depth genuinely goes unused.

Use Opus when you plan to argue with the answer. Anthropic's tutorial has an unusually specific line here, recommending Opus for "deep research and analysis you'll question, redirect, and build on as you go" and for "complex work you're doing in a live, back-and-forth session, where you want each answer sooner". That last clause is the reason Opus rather than Fable: Fable thinks for longer before it replies, which is fine for a job you leave running and irritating in a conversation.

One quirk worth knowing, because it means your choice is sometimes overridden. Anthropic's tutorial says that for biology and security topics, "Claude answers these topics with Opus even if you've picked Fable, so starting with Opus is simpler." Select the top tier, ask about those subjects, and you get the tier below regardless.

Claude Sonnet vs Haiku: This Is Where the Real Cliff Is

If the Sonnet-to-Opus step is smaller than the price list implies, the Haiku-to-Sonnet step is much larger.

Twenty points of index score separate them, against thirteen for Sonnet to Opus. On graduate-level science questions, Anthropic measured Haiku at 63% accuracy against Opus's 92%, at roughly a tenth of the cost per question, and noted that it "fell much further behind on long coding tasks". Its verdict is a single sentence and I would not improve on it: Haiku "fits high-volume work with checkable outputs, not long agentic loops."

Checkable is the operative word. Haiku is genuinely excellent at work where you would notice a mistake immediately. Pulling five figures out of a report you have read. Turning meeting notes into bullet points. Classifying a pile of emails. Rewriting a paragraph you will read anyway. In all of those, you are the verifier. What a deeper model buys you is being right in ways you would not have caught on your own, and if you were going to catch them anyway, that is not worth paying double for.

Where Haiku costs you is work you cannot check quickly, which tends to be the work that matters. That 200K context window is the practical limit most people hit first, not the accuracy score. And its knowledge cutoff sits eleven months behind Sonnet's and sixteen behind Fable's, so it is the tier most likely to state something confidently out of date.

The cost side does favour it enormously, and I do not want to undersell that. Finishing the same benchmark suite cost $0.21 on Haiku against $5.09 on Sonnet, a 24x difference, in a lineup where the Sonnet-to-Opus difference was 13%. Look at the whole ladder that way and the shape is not four evenly spaced rungs. It is one very large step from Haiku to everything else, then a fairly flat stretch across Sonnet, Opus and Fable. The real question is not which of four tiers to pick. It is whether this particular task is Haiku-shaped, and then which of the other three fits.

Anthropic suggests a test for exactly that, and it is a good one. Take a task you already know the answer to, such as summarising a report you have read yourself, run it on Haiku, then run it again on Sonnet in a new chat. "Compare where the answers differ, not how long they are." If Haiku caught what you would have caught, the task is Haiku-shaped and you can stop paying more for it.

The Second Dial Almost Nobody Mentions

Every comparison of these tiers treats the model picker as the only control. It is not, and Anthropic is direct about which one to reach for first: "Tuning effort is often a better lever than switching models."

Effort governs how much thinking, tool calling and self-verification the model does before answering. On claude.ai it sits in the model picker next to the model name. Sonnet, Opus and Fable all support it and all default to high. Haiku does not support it at all, which is one of those four spec-sheet rows.

The measured effect is larger than I expected. On Anthropic's research and knowledge-work benchmarks, the low setting gave up one to three points of accuracy for between a third and a half off the cost per task, and the medium setting matched the default's accuracy at 70% to 87% of the cost. The default bought nothing measurable over medium on any of the four benchmarks in that group. Lower effort is also faster, by about 40% on one of them.

Long, multi-step work is where effort genuinely buys accuracy rather than just spending money, and there the trade is real: on a long-horizon coding benchmark, Opus gave up about two points at medium for half the cost, and about eight points at low for a quarter of it.

The practical consequence for anyone on a subscription is that you have two dials, not one, and the second is cheaper to turn. If Sonnet is not getting there, the next move is not automatically Opus at full effort. Anthropic's own troubleshooting row says to restore effort first if you lowered it, and otherwise to "try the next tier up at low effort". On a rate-limited plan that is the frugal path: a more capable model, thinking less, often lands ahead of a weaker model thinking hard.

There is a warning attached to this too, which Anthropic states against its own more elaborate advice: in its internal measurements, "a multi-model configuration that looked cheaper than the default single model cost more than that same model at lower effort." Turn the simple dial before you build anything clever.

Price the Hardest Tenth of Your Work, Not the Typical Task

The last idea from Anthropic's cost guide is the one I would keep if I could keep only one, and it explains why so many people pick a tier, feel clever about the saving, and end up back where they started.

"Price the tail of your workload, not the median: compare models on the hardest tenth of your tasks, not the typical one. On the typical task every model looks similar and the cheapest looks best, but the bill is decided by the tasks the cheaper model fails, because a failed task still bills its tokens, then the retry, then whatever the failure costs downstream."

The measurement behind it is stark. On one 20-problem run, two problems carried 43% of the total spend.

Translate that off the API and onto a subscription and it reads even better, because on a plan you are not spending dollars, you are spending your allowance. A Haiku answer that misses something forces a re-ask on Sonnet, and you have now paid for both. The cheap tier is only cheap when it succeeds. Choose your default by looking at the tasks that went wrong last month, not the ones that went fine.

Which Tiers Your Plan Actually Reaches

None of this matters if the model you want is behind a paywall you are not on. Anthropic's pricing table settles it:

Tier Free Pro Max
Haiku Yes Yes Yes
Sonnet Yes Yes Yes
Opus No Yes Yes
Fable No Usage credits Up to 50% of weekly limits

Pro at $20 a month, or $17 if you pay for a year up front, is the line where Opus appears. Fable is the odd one. It is selectable on Pro but sits outside the plan's allowance, so every message runs on usage credits at standard API rates, billed separately from the subscription. The full mechanics of that meter are in the Claude limits post, and the plan prices are broken down in how much Claude costs.

Worth noticing: Anthropic's own tutorial says "Pro and Max add Opus, Fable, and more headroom", which reads as though both tiers come with the plan. The pricing table says otherwise for Fable. When two Anthropic pages disagree, the pricing table has been the reliable one.

Pick by Job

Stripping out the theory, this is where each tier earns its place for the kind of work most people actually do.

Haiku. Short questions with short answers. Pulling specific details out of text. Classifying or sorting. Quick summaries of things you have read. Anything where you would spot an error in the first sentence.

Sonnet. Writing and editing. Analysis that needs reasoning without needing depth. Working through a document. Multi-step tasks with clear checkpoints. If you want one default and no thought about it, this is the correct default.

Opus. Research you intend to interrogate. Reasoning where being wrong is expensive. Live back-and-forth sessions where you want depth without a long wait. Anything Sonnet has already tried and fumbled.

Fable. Long tasks with many connected steps, where you would rather describe the outcome than supervise each move. Dense source material: long documents, charts, technical diagrams. Work where accuracy matters more than the wait and more than the price.

The Version of This Question You Can Answer in One Thread

Every comparison above assumes you are choosing inside one vendor's lineup, on one vendor's meter. That assumption is doing more work than it looks.

The izzedo chat model picker open above the message box, with an Auto row at the top and GPT-5.6 Sol, GPT-5.6 Terra, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Grok 4.6, Perplexity Sonar Pro and Nano Banana 2 listed below it

izzedo chat puts all four Claude tiers, Haiku, Sonnet, Opus and Fable, in one picker alongside ChatGPT, Gemini, Grok, Perplexity Sonar, DeepSeek and around twenty more models, for $6/month, with a free plan that needs no card. Buying that lineup separately runs about $110/month across six accounts. Haiku is on the free plan; the paid plan that unlocks the other three tiers costs less than a third of what Pro does, and unlike Pro it does not leave the top tier outside the allowance.

Three things about that change the shape of this decision rather than just the price.

The effort dial is there too, per model, so the second lever in this post is a lever you actually have. So is a row at the top of the picker labelled Auto, which reads the message, works out what kind of task it is and how hard it is, then picks a model and an effort level for it. That is the entire subject of this post, handled per message, for the cases where you would rather not think about it.

Switching happens inside the thread, and the new model inherits everything above it. That turns Anthropic's own advice into something you can actually do: run the task on the cheap tier, and when the answer looks thin, re-ask in the same conversation on a deeper one without pasting anything. Sonnet drafts, Opus checks the reasoning, and neither of them starts from a blank page.

An izzedo chat conversation in which one model answers and a different model picks up the next reply after a mid-thread switch, with the model that generated each answer labelled underneath it

And the ladder stops being one company's ladder. Opus is not the only step up from Sonnet once Gemini and GPT are in the same picker, and a second opinion from a different vendor catches a different class of mistake than a second opinion from a bigger model in the same family. That is a one-minute habit and it is most of the argument for not routing everything through one vendor.

I should be straight about izzedo's own limits, having spent this whole post reading someone else's fine print. izzedo meters usage too, because these models cost real money every time they answer, and nothing here is unlimited. Fair use is metered on a rolling window of a few hours with a weekly pool behind it, and whatever the allowances are today sits on the pricing page rather than in a post that will age out of date. What is different is that the meter is one pool shared across every model, so moving a task from Opus down to Haiku, or across to Gemini, genuinely stretches what is left, and no tier is parked outside the plan the way Fable is on Pro.

The Bottom Line

The three-tier framing is out of date, and what replaces it has two dials rather than one.

Haiku is the only genuinely different animal in the lineup: a fifth of the context, no effort control, a knowledge cutoff eleven months behind the next tier up, the nearest retirement date, and a benchmark bill 24 times smaller. Use it for work you can check at a glance and it is a bargain. Use it for work you cannot and you will pay twice.

Between Sonnet, Opus and Fable, the per-token price list overstates the difference. Sonnet finished a full benchmark suite for 13% less than Opus while scoring thirteen points lower, because it wrote two and a half times as many tokens to get there. Keep Sonnet as the default, reach for Opus when you plan to argue with the answer, and try lowering effort before you decide a tier has failed you.

And if the honest answer to "which one" is "it depends on the task", the useful move is not to pick better. It is to stop picking once per subscription and start picking once per message, in a workspace where every tier is one click away.

Frequently asked questions

What is the difference between Claude Opus, Sonnet and Haiku?

Price, depth and speed, in that order of usefulness. Haiku is the cheapest and fastest, and it is the only one of the four with a 200K context window instead of 1M, no effort setting, and a knowledge cutoff more than a year older than the rest. Sonnet is Anthropic's suggested default for everyday work. Opus is the reasoning tier, and Anthropic's developer documentation now recommends starting there rather than with Sonnet. Above all three sits Fable, which most comparisons still leave out.

Is Claude Opus worth it over Sonnet?

More often than the price list suggests. Per token Opus costs two and a half times Sonnet, but Artificial Analysis measured the full cost of running its 10-evaluation index on each, and Sonnet came in at $5.09 against Opus at $5.86, a gap of 13%. The reason is verbosity: Sonnet wrote 370 million tokens to finish the same set of tasks, against 140 million from Opus. On short, easy work Sonnet is genuinely cheaper. On work that takes real effort, most of the saving disappears.

Which Claude model should I use for everyday work?

Anthropic's consumer tutorial says Sonnet, and that is a reasonable default for writing, analysis and research. Its developer documentation says to start with Opus and optimise down, because it is measuring cost per finished task rather than cost per token. On a subscription, where you are spending a rate-limit pool rather than a token bill, the tutorial's advice fits better: keep Sonnet as your default and move up to Opus for work you plan to argue with.

Is Claude Haiku good enough?

For short, checkable work, yes, and it is very cheap. On Anthropic's own measurement it answered GPQA Diamond questions at roughly a tenth of Opus's cost per question, but scored 63% against Opus's 92%. Anthropic's guidance is explicit that Haiku fits high-volume work with checkable outputs rather than long, multi-step tasks. The step from Haiku up to Sonnet is by far the biggest capability jump in the lineup.

Does Claude have a model above Opus?

Yes. Fable is Anthropic's most capable widely released model, priced at $10 per million input tokens and $50 per million output, twice Opus. It ranked first of 202 models on the Artificial Analysis intelligence index when I checked on 8 September 2026. On claude.ai it is not part of the Pro plan's allowance, so Pro subscribers run it on usage credits at API rates. There is also Mythos, which Anthropic lists with limited availability.

What is the effort setting in Claude, and does it matter?

It controls how much thinking, tool calling and self-checking the model does before it answers, and Anthropic says tuning it is often a better lever than switching models. On its research and knowledge-work benchmarks, the low setting gave up one to three points of accuracy for a third to a half off the cost, and medium matched the default's accuracy at 70% to 87% of the cost. It sits in the model picker on claude.ai. Haiku is the one tier that does not support it.

Ready to try multi-model AI workflows?

Access GPT, Claude, Gemini, Perplexity, and more — all in one place.

Start for Free

Related articles