Claude Haiku 5.5: Anthropic's small model, 75% cheaper to run
On 7 October 2026, Anthropic shipped Claude Haiku 5.5, its new fast little model. The mission: claw back the budget tier, where rivals had quietly undercut it on price.

On 7 October 2026, Anthropic unveiled Claude Haiku 5.5, billed as its fastest, smartest and cheapest small model yet. The target market is no secret: high-volume, cost-sensitive grunt work — summaries, compactions, database queries, classification — plus playing sidekick to Opus 5.5 and Sonnet 5.5 on coding jobs. And because it is the quickest thing in the house, Anthropic is also nudging it toward live customer support and driving a browser.
Playing catch-up on price
At the budget end, price is the whole ballgame. The outgoing Haiku 4.5, released almost a year ago, charged $1 per million input tokens and $5 per million out — ten times the sticker on OpenAI’s GPT-6 Luna, which landed last month, as Simon Willison points out. Haiku 5.5 matches Luna penny for penny: $0.10 in, $0.50 out. Anthropic, for its part, says the thing costs on average 75% less to run than Haiku 4.5.
Two asterisks, and they earn their keep. That price only holds up to 100,000 tokens; past that it jumps fivefold, to $0.50 / $2.50. Luna doesn’t shift gears until 272,000 tokens, and it shifts more gently ($0.20 / $0.75). So under 100,000 tokens Haiku claims better scores for the same money; above it, Luna quietly becomes the better deal again.
The tokenizer that bills twice
Willison flags another wrinkle that never made it onto the price sheet: Haiku 5.5 ships a new, stingier tokenizer. His own token counter shows the same long prompt eating roughly 1.25 times more tokens than it did on Haiku 4.5. Which means part of that headline discount quietly evaporates, because the identical request now weighs more on the meter. The price per token drops; the token count climbs to meet it.
An effort dial, at last
On the brains front, a genuine first: this is the opening Haiku model to offer an adjustable effort setting, the kind the flagship models already have. Five levels — low, medium, high, xhigh, max — let you trade cost against quality. One quirk: you can’t switch the thinking off entirely, and the dial starts at medium. Anthropic publishes curves across three benchmarks measured at every level — OSWorld (computer use), GDPval-AA (knowledge work) and Humanity’s Last Exam (reasoning).
Among the early adopters, Asana reports that on its own agent test suite, task-completion latency fell by more than 30%, with inference running up to 2.5 times faster per agent turn than the model it had been using.
A little something for subscribers
Around the launch, Anthropic is halving the price of Sonnet 5.5 cache reads — roughly 20% cheaper on most agentic workloads. And, more enticingly, a monthly API credit arrives for Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team — which happens to be exactly what the plan costs. One catch Willison notes: the credits don’t roll over from one month to the next. Spend it or lose it.
So Anthropic has pulled level with OpenAI’s budget pricing — as long as you stay the right side of 100,000 tokens, and don’t look too hard at the little wheel spinning on the token meter.
Sources (2)
- Claude Haiku 5.5anthropic.com
- Claude Haiku 5.5simonwillison.net
Written with AI assistance from the sources cited above, then reviewed and approved before publication by Sébastien Soulier.


