Anthropic shipped a new small model on October 7, 2026, and the headline number is the price tag, not the benchmark chart. Claude Haiku 5.5 is now live on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, and Anthropic says it costs around 75% less to run on average than its predecessor, Claude Haiku 4.5. For a company that built its brand on frontier reasoning models, leading with a cost cut on the cheap tier says something about where the real AI spending battle is happening right now.
Anthropic calls Haiku 5.5 its “cheapest, fastest, and most capable small model” to date, aimed squarely at the unglamorous but expensive work that sits behind most production AI systems: summarizing documents, compacting context windows, running database queries, classifying support tickets, and powering voice agents and browser-use bots that fire off thousands of small requests an hour. None of that work needs a flagship model. All of it needs to be cheap at scale, and that’s the gap Anthropic is trying to close.
A new Haiku lands in the middle of an AI price war
Haiku 5.5 did not arrive in isolation. Anthropic shipped Claude Opus 5.5 and Claude Sonnet 5.5 within six days of each other earlier this cycle, and industry watchers had already flagged Haiku as the next piece of the 5.5 generation to land. Today’s release closes that loop and gives Anthropic a full three-tier lineup running the same version number for the first time this generation.
The timing matters. Model vendors have spent 2026 undercutting each other on inference pricing almost as aggressively as they’ve competed on benchmark scores. Google cut its Gemini 4 Argon pricing to $2 per million input tokens and $10 per million output tokens earlier this year, and Mistral has pushed its own large-model economics with Large 4, a 675-billion-parameter model with 41 billion active parameters, built to compete on cost-per-token at scale. Anthropic’s move on Haiku 5.5 is the small-model version of that same fight, aimed at the high-volume, low-margin workloads where every fraction of a cent per call adds up fast.
What Claude Haiku 5.5 actually changes
The model carries the identifier claude-haiku-5-5 in the API, a straightforward naming choice that lets existing integrations swap in the new model without restructuring request logic. Anthropic’s own description keeps things simple: “Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks.” That’s not a model built to win leaderboard arguments. It’s built to sit quietly inside a pipeline and process a very large number of small, repetitive requests without blowing up a cloud bill.
That positioning lines up with how Anthropic has talked about its three current model tiers. Opus handles the hardest reasoning and agentic coding work, Sonnet sits in the middle as the general-purpose workhorse, and Haiku now exists specifically for throughput: tasks measured in thousands of calls per hour rather than depth of reasoning per call. The Opus 5.5 release brought a 1-million-token context window for coding tools, which tells you where Anthropic wants its top-tier model spending its compute. Haiku 5.5 tells you the opposite: spend as little compute as possible, as fast as possible, as often as needed.
The pricing cut, tier by tier
Anthropic published two pricing tiers for Haiku 5.5, split by prompt length. The split itself is not new to Anthropic’s pricing model, but the gap between the new rates and the old Haiku 4.5 rates is the story here.
Short prompts get the deepest discount
For prompts up to 100,000 tokens, which covers the overwhelming majority of classification, summarization, and short customer-support exchanges, Haiku 5.5 runs at $0.10 per million input tokens and $0.50 per million output tokens. Cache reads drop to $0.01 per million tokens and cache writes sit at $0.125 per million. That’s the tier where Anthropic’s reported savings are steepest, and it’s also the tier where most high-volume workloads actually live.
Long-context requests still get cheaper, just less dramatically
Once a prompt crosses 100,000 tokens, pricing shifts to $0.50 per million input tokens and $2.50 per million output tokens, with cache reads at $0.05 per million and cache writes at $0.625 per million. The discount narrows at this tier, which makes sense: long-context requests consume more compute regardless of which model processes them, so the savings from a smaller, faster model matter less once the context window itself becomes the dominant cost driver.
| Pricing tier | Input (per million tokens) | Output (per million tokens) | Cache reads | Cache writes |
|---|---|---|---|---|
| Up to 100,000 tokens | $0.10 | $0.50 | $0.01 | $0.125 |
| Over 100,000 tokens | $0.50 | $2.50 | $0.05 | $0.625 |
Here’s what that looks like as an API call. Developers migrating from Haiku 4.5 mainly need to change the model string, since Anthropic kept the request shape consistent across the Haiku line:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2026-01-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 512,
"messages": [
{"role": "user", "content": "Classify this support ticket by urgency and topic."}
]
}'
Where Claude Haiku 5.5 runs
Anthropic says Haiku 5.5 is available immediately on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure. That four-cloud spread matters more for a cost-tier model than it would for a flagship one, because teams running high-volume pipelines tend to be locked into whatever cloud already hosts their data warehouse or support stack. A model that only shipped on one cloud would cut off a meaningful slice of the exact customers Anthropic is trying to win with this release.
| Platform | Status |
|---|---|
| Claude Platform (Anthropic direct API) | Available at launch |
| Amazon Web Services | Available at launch |
| Google Cloud | Available at launch |
| Microsoft Azure | Available at launch |
Adjustable effort: a new dial for small models
Haiku 5.5 is Anthropic’s first Haiku-class model to ship with an adjustable effort setting, a control that lets developers trade latency and cost against reasoning depth on a per-request basis. Previously, effort controls of this kind were mostly a Sonnet and Opus feature, used when a task needed the model to think longer before answering. Bringing that same dial down to the small-model tier gives teams a way to squeeze extra accuracy out of Haiku for the subset of requests that need it, without paying Sonnet or Opus prices for the rest of the traffic.
In practice, that means a classification pipeline could run most tickets at low effort for speed and cost, then bump a smaller slice of ambiguous cases up to a higher effort setting rather than routing them to a larger, more expensive model entirely. It’s a small architectural change, but it pushes more of the cost-versus-accuracy decision down into the cheapest model in the lineup, which is exactly where high-volume teams want that decision to live.
Who Anthropic built this model for
Anthropic’s own framing lists the target workloads plainly: summaries, compactions, database queries, classification, live customer support, browser use, voice agents, and in-app assistants. Every one of those categories shares a trait that explains the pricing strategy. They run at volume, they tolerate a smaller model’s ceiling on reasoning, and they’re exactly the kind of workload that breaks a budget when the per-call cost doesn’t scale down as usage scales up.
Developers already building agentic systems with Anthropic’s tooling, including teams working through the Claude Code workflow, are a natural early audience here too. Agentic coding tools lean on a mix of model tiers: a capable model drives the core reasoning, while a cheaper model handles side tasks like summarizing a long file, compacting conversation history, or triaging which files are relevant to a query. Haiku 5.5’s cost structure makes that kind of multi-model architecture noticeably cheaper to run at scale.
Claude Haiku 5.5 vs. Claude Haiku 4.5, by the numbers
Anthropic’s headline figure is a 75% average cost reduction compared with Haiku 4.5. That’s a blended number across both pricing tiers and real-world traffic mixes, so it reads differently depending on which tier a given workload falls into. Short-prompt traffic, which both Anthropic and multiple reporting outlets describe as the bulk of typical small-model usage, sees the steepest cut. Longer-context traffic sees a smaller, though still meaningful, reduction.
| Metric | Claude Haiku 4.5 | Claude Haiku 5.5 |
|---|---|---|
| Average cost to run (reported) | Baseline | Around 75% lower |
| Adjustable effort setting | Not available | First Haiku model to ship with it |
| Cloud availability | Claude Platform, AWS, Google Cloud, Azure | Claude Platform, AWS, Google Cloud, Azure |
| Positioning | Cheapest small model at prior release | “Cheapest, fastest, and most capable small model” per Anthropic |
What’s notable is what Anthropic didn’t lead with here: a benchmark score. Unite.AI’s coverage of the launch pointed to a system card and reported benchmark results that show gains over Haiku 4.5 on agentic and tool-use tasks, which suggests Anthropic isn’t treating this purely as a cost-cutting release. But the public announcement itself centers on price, not score, a framing choice that says the company expects price to be the deciding factor for most buyers in this tier.
How Haiku 5.5 fits inside Anthropic’s model lineup
With Opus 5.5, Sonnet 5.5, and now Haiku 5.5 all shipped, Anthropic has a matched-generation lineup across all three tiers for the first time in this cycle. That’s a different posture than the staggered releases that defined earlier generations, where a new Opus might ship months before its matching Sonnet or Haiku counterpart. Running all three on the same version number simplifies the pitch to enterprise buyers: pick a tier based on task complexity and budget, not on which generation happens to be current for that specific tier.
| Model | Role in lineup | Primary use case |
|---|---|---|
| Claude Opus 5.5 | Top-tier reasoning and coding | Complex agentic work, long-context coding tasks |
| Claude Sonnet 5.5 | General-purpose mid-tier | Balanced reasoning and cost for everyday production use |
| Claude Haiku 5.5 | High-volume, low-cost tier | Summaries, classification, support agents, voice and browser agents |
Six days, three models: the small-model price war in context
Anthropic’s pace this cycle has been unusually tight. Shipping Opus and Sonnet six days apart, then following with Haiku in the same broader window, compresses what used to be a multi-month release cadence into a single sprint. That pace shows up elsewhere in Anthropic’s recent disclosures too. The company’s IPO filing devoted roughly 80 pages to AI risk, a sign that Anthropic is trying to move fast on releases while simultaneously building a much larger paper trail around safety and governance than it has in prior cycles. Shipping a cost-focused small model fits a company trying to win market share quickly while keeping its risk disclosures thorough enough to satisfy public-market scrutiny.
The broader market has been circling the same question for most of 2026: can a lab win the inference-cost fight and the capability fight at the same time? Mistral’s bet with Large 4’s 41-billion active-parameter design is one answer, using sparse activation to keep large-model costs down. Anthropic’s answer with Haiku 5.5 is simpler: build a genuinely small model, keep its scope narrow, and let the cost advantage come from the model’s size rather than from architectural tricks layered on top of a bigger one.
Market impact: what cheaper inference does to AI budgets
A 75% cost cut on a small model doesn’t sound as dramatic as a new frontier model topping a leaderboard, but for teams running inference at scale it can matter more to the bottom line. High-volume workloads are exactly where API costs compound fastest, because the per-call price gets multiplied by call counts that run into the millions. A support team routing every inbound ticket through a classification model, or an agent framework compacting conversation history on every turn, feels a price cut like this immediately in its monthly cloud bill.
Enterprises running high-volume pipelines
For enterprise teams, the practical effect of a cut this size is less about saving money on existing workloads and more about which workloads become viable to automate at all. A classification task that was borderline too expensive to run on every record now clears the budget bar. A voice agent that needed to limit how often it called out to a model for intent classification can afford to call more often. Cheaper small-model inference tends to expand the footprint of AI automation rather than simply shrinking existing bills, because teams redeploy the savings into running the model more often, not just running it more cheaply.
Developer and early-adopter reaction
Coverage from 9to5Mac framed the release around the same cost-reduction figure Anthropic published, describing Haiku 5.5 as “the cheapest, fastest, and most capable small model” Anthropic has shipped, while noting that Sonnet 5.5’s cache-read pricing also dropped as part of the same pricing refresh. Startup Fortune’s write-up centered on the same headline figure, reporting the launch as part of a broader pattern of price reductions reaching as high as 90% depending on prompt length. Yahoo Finance’s coverage placed the release inside the wider competitive context, describing it as part of an intensifying AI pricing war rather than an isolated product update.
That consistency across outlets, reporting the same per-tier pricing numbers rather than conflicting figures, suggests Anthropic’s announcement was specific enough that there wasn’t much room for ambiguity. The 90% figure some outlets cite refers specifically to the under-100,000-token tier, while the 75% figure is the blended average Anthropic uses for overall workload cost. Both numbers are accurate descriptions of the same pricing table, just measuring different slices of it.
Competitive response: what rivals do next
Anthropic isn’t the only lab leaning on small-model pricing as a competitive lever this year. Google’s Gemini line has gone through its own pricing and tier restructuring with Gemini 4 Argon, and OpenAI has run its own cost-tiering strategy across its model lineup throughout 2026. The pattern across all three labs looks similar: ship a flagship model to win the capability argument, then follow with cheaper, smaller models to win the volume argument, since most production traffic by call count runs through the cheap tier rather than the flagship one.
The open question is how quickly rivals respond specifically to Haiku 5.5’s pricing table. Price wars in this market tend to move fast once one lab publishes a new low. If Google or OpenAI match or undercut Anthropic’s new rates on their own small-model tiers within the next few weeks, it would confirm that small-model pricing has become as competitively reactive as flagship-model benchmark racing already is.
Risks and open questions
A model built for high-volume, low-oversight tasks like classification and voice-agent routing raises a different risk profile than a flagship reasoning model. Errors in a cheap model running millions of calls a day don’t show up as single dramatic failures. They show up as a steady background rate of misclassified tickets, mis-routed calls, or subtly wrong summaries, the kind of problem that’s easy to miss until it’s been running at scale for weeks. Anthropic hasn’t published detailed error-rate comparisons against Haiku 4.5 in its public pricing announcement, so teams evaluating the migration will need to run their own accuracy testing against their specific workloads rather than assuming the price cut comes with zero tradeoff.
There’s also a structural question about what happens to the adjustable effort setting at scale. Letting developers dial effort up for harder requests is useful, but it also means two teams running what looks like the same model can end up with very different cost and accuracy profiles depending on how aggressively they tune that dial. That’s a new variable for teams to manage that didn’t exist in the Haiku 4.5 era.
What to watch next
- Expect rival labs to respond to Haiku 5.5’s pricing within weeks rather than months, given how reactive small-model pricing has been across 2026.
- Expect Anthropic to extend the adjustable effort setting to Sonnet 5.5 if early Haiku 5.5 adoption shows it meaningfully improves the cost-accuracy tradeoff for developers.
- Expect enterprise teams already running Haiku 4.5 in production to migrate quickly, since the API only requires a model-string change rather than a structural rework.
- Expect growth in agentic pipelines that mix model tiers, using Haiku 5.5 for high-volume side tasks like summarization and compaction while a larger model handles core reasoning.
- Expect continued scrutiny of error rates at scale, since Anthropic’s public announcement emphasizes cost over detailed accuracy benchmarking against Haiku 4.5.
Frequently asked questions
What is Claude Haiku 5.5?
Claude Haiku 5.5 is Anthropic’s newest small language model, released October 7, 2026. Anthropic describes it as the cheapest, fastest, and most capable small model it has shipped, built for high-volume, cost-sensitive tasks rather than complex reasoning work.
How much does Claude Haiku 5.5 cost compared to Haiku 4.5?
Anthropic says Haiku 5.5 costs around 75% less to run on average than Haiku 4.5. For prompts up to 100,000 tokens, it’s priced at $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens, pricing rises to $0.50 input and $2.50 output per million tokens.
Where can developers access Claude Haiku 5.5?
Anthropic says the model is available on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, all starting on launch day.
What is the model identifier for Claude Haiku 5.5 in the API?
The model identifier is claude-haiku-5-5, which developers can use in place of the Haiku 4.5 identifier in existing API calls.
What is the adjustable effort setting in Claude Haiku 5.5?
It’s a control that lets developers trade latency and cost against reasoning depth per request. Anthropic says Haiku 5.5 is the first model in the Haiku line to support it, giving high-volume applications a way to selectively use more compute on harder requests.
What tasks is Claude Haiku 5.5 designed for?
Anthropic lists summaries, compactions, database queries, classification, live customer support, browser use, voice agents, and in-app assistants as the intended uses, all workloads characterized by high call volume and lower per-call reasoning requirements.
Is Claude Haiku 5.5 part of a larger model family release?
Yes. Anthropic also named Claude Opus 5.5 and Claude Sonnet 5.5 as related models in the same generation, giving the company a matched three-tier lineup across reasoning, general-purpose, and high-volume use cases.
Does Claude Haiku 5.5 match OpenAI or Google’s small-model pricing?
Anthropic’s own materials don’t make a direct pricing comparison to a specific competing model. What’s clear from the announcement is that Haiku 5.5’s pricing lands well below Haiku 4.5’s rates, in a market where Google and other labs have also been cutting small-model and mid-tier pricing throughout 2026.




