What Is Claude Haiku 5.5? New Features, API Pricing and How It Compares With Haiku 4.5
Claude Haiku 5.5 is Anthropic’s new small AI model, released on October 7, 2026, for work that needs quick responses at low cost. Anthropic says it is its fastest and cheapest small model yet and costs about 75% less to run than Claude Haiku 4.5 on average. For developers, its lowest API rates are $0.10 per million input tokens and $0.50 per million output tokens, though longer prompts cost more.
What is new in Claude Haiku 5.5?
Haiku 5.5 is aimed at frequent, narrowly scoped jobs rather than serving as Anthropic’s most powerful model for every task. The company identifies summaries, condensing conversations, database queries and classification as suitable uses. It also points to information extraction and routing requests to the right place. Speed-sensitive uses include live customer support and browser-based tasks.
For coding workflows, Anthropic suggests using Haiku 5.5 as a subagent: a larger model can handle the main problem while Haiku completes smaller supporting steps. Anthropic says Sonnet 5.5 and Opus 5.5 remain better choices for complex, multi-step coding work.
Compared with Haiku 4.5, the new model adds adaptive thinking and an adjustable effort setting. Developers can use that setting to balance response quality against speed and cost for a particular task. Its context window also grows from 200,000 to 1 million tokens, while its maximum output rises from 64,000 to 128,000 tokens. Those larger limits provide room for longer inputs and responses, but do not mean every request needs to use them.
Anthropic reports improvements over Haiku 4.5 in its evaluations of coding, computer use and knowledge work. Those are company-reported results, not a guarantee that the newer model will outperform its predecessor in every application. Teams considering a switch should test the tasks and response times that matter to them.
How much does Claude Haiku 5.5 cost?
For API requests with prompts of up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. For prompts over 100,000 tokens, the rates rise to $0.50 for input and $2.50 for output per million tokens. These are developer API rates, not the price of a Claude subscription.
Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens. On a per-token basis, the new model’s rates are therefore 90% lower for prompts of up to 100,000 tokens and 50% lower for longer prompts. Anthropic says roughly 90% of requests to Haiku 4.5 fell into the shorter-prompt group.
Why does Anthropic say the average cost to run Haiku 5.5 is about 75% lower, rather than 90%? The average accounts for both prompt lengths and a change in how text is counted. Anthropic’s developer documentation says the same input text produces approximately 30% more tokens with Haiku 5.5’s newer tokenizer than with Haiku 4.5, although the difference varies by content. A lower price per token does not always translate into an identical reduction on a finished job.
Developers using prompt caching have separate rates. A cache read costs $0.01 per million tokens for prompts up to 100,000 tokens, or $0.05 above that threshold. Anthropic also lists a 50% discount on input and output tokens through its Batch API. Actual bills will depend on the mix of inputs, outputs, cached material and prompt lengths.
Where is it available?
Haiku 5.5 is available through Anthropic’s Claude API under the model ID claude-haiku-5-5, as well as through Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic says it is available across its platforms.
Switching an existing application from Haiku 4.5 may take more than changing the model name. Anthropic advises developers to recount tokens and review output limits; some older thinking and sampling settings are not compatible with Haiku 5.5. For most readers, the central distinction is simpler: Haiku 5.5 is built to make everyday, repeated AI tasks faster and cheaper, while the largest Claude models remain the stronger option for demanding work.

