Claude Haiku 5.5 launched on October 7, 2026, as Anthropic’s new small model for high-volume AI work. Anthropic says its average running cost is about 75% below Haiku 4.5. That is the vendor’s workload-based estimate, not a universal bill reduction: list prices vary when prompts exceed 100,000 tokens, and the new tokenizer can count the same text differently.
The clearest change is the short-prompt list price: US$0.10 per million input tokens and US$0.50 per million output tokens, versus US$1 and US$5 for Haiku 4.5. SiliconANGLE, Decrypt, and DataCamp reported the launch independently. Anthropic also introduced adjustable effort for Haiku, a one-million-token context window and lower Sonnet 5.5 cache-read rates. Broad benchmark gains remain Anthropic’s claims unless independently reproduced.
Key takeaways: Haiku 5.5 is available through Anthropic and major cloud platforms. The lowest API rates apply to prompts of no more than 100,000 tokens; longer prompts cost five times as much per token. The 75% average-savings figure depends on workload mix and tokenization. Teams should compare costs and task quality on their own data before moving production traffic.
What Is Claude Haiku 5.5?
Claude Haiku 5.5 is Anthropic’s latest small AI model, built specifically for workloads where speed, cost and scale matter.
Anthropic describes it as its cheapest, fastest and most capable small model to date. Unlike larger models such as Opus 5.5, Haiku 5.5 is intended to handle large volumes of relatively focused tasks rather than serve as the primary model for every complex AI workflow.
The model can handle text and image inputs and generate text outputs. It has a June 2026 knowledge cutoff, a 1-million-token context window and can generate up to 128,000 output tokens.
Claude Haiku 5.5 at a Glance
| Feature | Claude Haiku 5.5 |
|---|---|
| Launch date | October 7, 2026 |
| Context window | 1 million tokens |
| Maximum output | 128,000 tokens |
| Input pricing | From $0.10 per million tokens |
| Output pricing | From $0.50 per million tokens |
| Reasoning | Adaptive thinking |
| Effort control | Yes |
| Relative speed | Anthropic’s fastest model |
| Knowledge cutoff | June 2026 |
| Main focus | High-volume, latency-sensitive workloads |
These specifications make Haiku 5.5 particularly suited to applications that need to make many AI calls without incurring the costs associated with larger frontier models.
Claude Haiku 5.5 Is Much Cheaper
The biggest headline from the launch is pricing.
For prompts up to 100,000 tokens, Anthropic charges $0.10 per million input tokens and $0.50 per million output tokens.
For prompts exceeding 100,000 tokens, the rates increase to $0.50 per million input tokens and $2.50 per million output tokens.
That represents a substantial reduction from Claude Haiku 4.5.
Anthropic says Haiku 5.5 costs around 75% less to run on average than its predecessor. For many shorter workloads, the difference is even larger because the new list price is roughly one-tenth of Haiku 4.5’s input and output rates.
Claude Haiku 5.5 Pricing
| Model | Input, up to 100K tokens | Output, up to 100K tokens |
|---|---|---|
| Claude Haiku 5.5 | $0.10 / 1M tokens | $0.50 / 1M tokens |
| Claude Haiku 5.5, above 100K | $0.50 / 1M tokens | $2.50 / 1M tokens |
| Claude Sonnet 5.5 | $2 / 1M tokens | $10 / 1M tokens |
| Claude Opus 5.5 | $4 / 1M tokens | $20 / 1M tokens |
The pricing makes Haiku 5.5 particularly attractive for applications that may generate millions of AI requests.
Built for High-Volume AI Work
Anthropic is targeting Haiku 5.5 at workloads where AI needs to operate continuously and at large scale.
Potential applications include:
- Customer support
- Document summarisation
- Classification
- Data extraction
- Database queries
- Routing
- Browser automation
- Voice agents
- In-app assistants
- AI subagents
- Context compaction
For these applications, even a small difference in the cost of an individual AI request can become significant when multiplied across millions of calls. Anthropic is therefore positioning Haiku 5.5 as an infrastructure model rather than simply another chatbot.
Why Lower Costs Matter
Consider a company processing millions of customer-support interactions every month.
If each request becomes cheaper while maintaining acceptable quality, the company can potentially:
- Process more requests.
- Reduce its AI infrastructure bill.
- Add AI to additional workflows.
- Run more subagents simultaneously.
- Use AI for tasks that were previously too expensive.
This is one reason small models have become increasingly important in the AI industry.
Haiku 5.5 Adds Adjustable Thinking
Another important addition is adaptive thinking with an effort control.
Haiku 5.5 is the first model in Anthropic’s Haiku family to include this capability. The default effort level is medium, but developers can adjust the amount of reasoning used by the model.
This gives developers greater control over the trade-off between speed, cost and reasoning depth.
A simple classification request may require very little reasoning.
A more complicated subagent task may benefit from spending additional computation before producing an answer.
One Model, Different Levels of Effort
The adjustable approach means developers do not necessarily have to switch models every time a task becomes slightly more difficult.
Instead, they can potentially increase or decrease the model’s effort level depending on the task.
This could be useful in automated systems where thousands of different requests arrive with varying levels of complexity.
1 Million-Token Context Window
Haiku 5.5 also comes with a 1-million-token context window, matching the large context capabilities of Anthropic’s higher-end models. It can generate up to 128,000 tokens in a single response.
A large context window allows an AI system to process much more information in one interaction.
For example, a business application could provide a model with:
- Large documents.
- Multiple reports.
- Long software repositories.
- Extensive customer histories.
- Large amounts of structured information.
The model can then reason over that information without necessarily requiring developers to break it into many smaller requests.
Haiku 5.5 Is Designed to Work With Bigger Models
Anthropic does not necessarily expect Haiku 5.5 to replace its larger models.
Instead, the company specifically highlights its ability to work alongside Opus 5.5 and Sonnet 5.5 as a subagent.
This creates a hierarchy of AI models.
A powerful model such as Opus 5.5 could manage a complicated task, while Haiku 5.5 handles smaller jobs required along the way.
For example, an advanced AI agent could ask Haiku to:
- Extract information from a document.
- Search through a large dataset.
- Classify results.
- Summarise findings.
- Perform repetitive coding tasks.
- Check individual pieces of information.
The larger model can then use those outputs to complete the broader task.
AI Agents Could Become Cheaper
This could be particularly important as AI agents become more common.
An agent may need to make dozens or hundreds of model calls during a single workflow.
Using a large model for every step could become expensive.
A cheaper model such as Haiku 5.5 can potentially perform many of the simpler steps while reserving expensive models for the hardest decisions.
That could make large-scale AI agents more economically practical.

Haiku 5.5 Targets Coding and Computer Use
Anthropic reports stronger coding and agent results than Haiku 4.5 in its own evaluations. DataCamp also published a hands-on test, but that limited test does not establish a broad ranking across all tasks.
The company particularly highlights browser use and subagent coding as applications where the model’s speed can be useful.
This is important because AI coding systems often perform many smaller operations.
An AI coding agent might need to inspect a file, search for a function, classify an error, execute a command or summarise test results before the main model makes a larger decision.
Haiku 5.5 can potentially handle some of those supporting tasks at much lower cost.
Anthropic Is Cutting Costs Across Its Model Lineup
The Haiku 5.5 launch is accompanied by another pricing change.
Anthropic is cutting the price of Claude Sonnet 5.5 cache reads by half.
The company says the change makes Sonnet 5.5 around 20% cheaper on most agentic workloads.
Anthropic is also introducing monthly API credits for Claude Max and Team subscribers who want to build their own applications and agents on the Claude Platform.
These moves indicate that Anthropic is focusing not only on model performance but also on the economics of deploying AI at scale.
The AI Industry Is Entering a Price Competition
The Haiku 5.5 launch comes during an increasingly competitive period for AI model pricing.
Companies are trying to deliver stronger models while simultaneously reducing the cost of running them.
That matters because AI adoption is increasingly moving from individual chatbot conversations to large automated workloads.
A company might make thousands or millions of API calls every day.
At that scale, model pricing can directly influence whether an AI application is commercially viable.
Anthropic’s lower Haiku pricing therefore targets an important part of the market where small improvements in efficiency can have a large financial impact.
Where Haiku 5.5 Fits in Anthropic’s Lineup
With Haiku 5.5, Anthropic now has a three-tier Claude 5.5 family.
| Model | Primary positioning | Relative cost |
|---|---|---|
| Claude Opus 5.5 | Most demanding reasoning and complex agentic work | Highest |
| Claude Sonnet 5.5 | General-purpose performance and coding | Mid-range |
| Claude Haiku 5.5 | High-volume, fast and cost-sensitive workloads | Lowest |
The structure gives developers the option of matching the model to the complexity of a task.
Not every AI request needs the most powerful available model.
Availability
Claude Haiku 5.5 is available from launch through Anthropic’s Claude API and Claude Platform, as well as Amazon Bedrock, Google Cloud and Microsoft Foundry.
The model identifier for developers using the Claude API is claude-haiku-5-5.
Anthropic has also set a model retirement date no earlier than October 7, 2027, according to its platform documentation.
The Bigger Picture
Claude Haiku 5.5 shows how the AI competition is increasingly moving beyond the question of which company has the most powerful model. For businesses, the cost and speed of running AI can be just as important as benchmark performance.
Anthropic is targeting a massive market of automated AI tasks where companies need models to process information continuously. Anthropic’s claimed 75% average reduction in running costs, combined with its reported faster responses and a 1-million-token context window, could make Haiku 5.5 attractive for developers building large-scale AI agents and applications.
The model also demonstrates the growing importance of multi-model AI systems. A powerful model such as Opus or Sonnet can manage a complex workflow while a cheaper Haiku model handles repetitive supporting tasks. If that architecture becomes widespread, AI companies will increasingly compete on the economics of entire agent systems rather than just individual model benchmarks.
Looking Ahead
Anthropic is likely to continue improving the balance between AI capability, speed and cost as developers deploy increasingly complex agents. Haiku 5.5 gives the company a lower-cost model for high-volume workloads while Opus 5.5 and Sonnet 5.5 handle more demanding tasks, creating a more complete model portfolio.
For businesses and developers, the most important question will be whether Haiku 5.5 can maintain enough quality while operating at dramatically lower cost. If it can, the model could help make AI-powered customer support, browser agents, coding assistants and automated data processing substantially cheaper to operate, accelerating the shift from experimental AI tools to large-scale production systems.
How to assess the cost claim
The 75% figure is Anthropic’s estimate for average operating cost across its own request mix. The company says 90% of Haiku 4.5 requests used prompts below the 100,000-token threshold, where Haiku 5.5’s per-token list prices are one-tenth of the old rates. Above that threshold, new prices are half the old rates. Anthropic also says the newer tokenizer may count more tokens for the same work. Those factors help explain why a nominal 90% short-prompt rate cut does not mean a universal 90% invoice cut.
As an illustration, one million short-tier input tokens plus one million short-tier output tokens cost $0.60 at Haiku 5.5 list rates, versus $6.00 at Haiku 4.5 list rates. This is arithmetic, not a production-cost forecast. Caching, tool calls, long contexts, output length and retries can materially change a bill. Indian buyers paying in rupees should also consider currency conversion and applicable taxes.
Teams should benchmark quality, latency and cost per successfully completed task on representative requests. A cheaper model can cost more per successful workflow if it needs extra retries or a larger model to correct its output. This matters for multi-step agents. Our related coverage of Windows execution containers for AI agents addresses a different operational issue: constraining what an agent can access. Anthropic’s startup credits programme can reduce an initial bill, but credits do not alter the underlying per-token rate after they expire.
Frequently asked questions
When did Claude Haiku 5.5 launch?
Anthropic announced the model on October 7, 2026. Its release notes and model documentation list the same date. The model is available through Anthropic’s API and cloud platforms named in the launch post.
Is it always 75% cheaper?
No. That is Anthropic’s average running-cost estimate. The short-prompt per-token list rate is 90% below Haiku 4.5’s, while the long-prompt rate is 50% below it. Bills depend on the workload and tokenizer.
Are the performance benchmarks independently verified?
Anthropic published broad comparisons in its launch announcement. DataCamp tested the model separately, and Decrypt reported that its simple logic test produced a wrong answer. Neither limited test establishes a universal ranking, so users should test their own tasks.
Sources and verification
Primary sources: Anthropic’s October 7 announcement and pricing table, model documentation, and release notes. Independent original reports: SiliconANGLE, Decrypt, and DataCamp’s hands-on analysis. Prices are launch-date US-dollar list prices. Vendor claims are identified as Anthropic’s and the example calculation is illustrative.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



