TL;DR
Anthropic launched Claude Haiku 5.5 on October 7, the third and final model in its Claude 5.5 family. For prompts under 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens — a 90% cut from Haiku 4.5 — and Anthropic reports big capability jumps, like 72.4% on the OSWorld computer-use benchmark versus 15.7% for its predecessor. Sonnet 5.5 cache reads were also halved, and Max and Team subscribers get new monthly API credits.
What launched
Claude Haiku 5.5 arrived October 7, 2026, completing the Claude 5.5 family after September's Opus 5.5 and Sonnet 5.5 releases. Anthropic describes it as its cheapest, fastest, and most capable small model — built for high-volume, cost-sensitive jobs like summaries, context compaction, database queries, and classification, and as a subagent alongside the bigger models on coding work. It is available now under the model ID claude-haiku-5-5 on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, plus Claude.ai and Claude Code.
One practical note: this is the first Haiku-class model with an adjustable effort setting, so developers can dial a request toward lower cost or higher intelligence without switching models. Anthropic also updated its Claude Python and TypeScript SDKs to add computer use and browser use in beta, and says Haiku 5.5's speed, capability, and price make it well-suited to those tasks.
The price cut — and the fine print
The headline is real: for prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, down from Haiku 4.5's $1.00 and $5.00 — a 90% reduction. Above 100,000 tokens, the rates jump to $0.50 and $2.50, five times the lower tier but still half of Haiku 4.5. Anthropic says roughly 90% of requests to the old Haiku fell inside the cheaper tier.
The company's own math puts average running costs about 75% below Haiku 4.5, not 90% — the calculation accounts for the tier mix and a new tokenizer that uses somewhat more tokens per task, a detail VentureBeat highlighted. For comparison, OpenAI's GPT-6 Luna matches Haiku's lower-tier rates exactly, while Google's Gemini 3.5 Flash-Lite sits at $0.30 in and $2.50 out. Alongside the launch, Anthropic halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens — about a 20% saving on typical agentic work — and is rolling out monthly API credits: $100 for Max 5x subscribers, $200 for Max 20x, and up to $500 pooled for Team.
The benchmarks (reported by Anthropic)
Treat these as vendor numbers — VentureBeat explicitly flags them as not independently verified — but the jumps are large. On OSWorld 2.1's offline computer-use subset, Haiku 5.5 scored 72.4% versus 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna. On Terminal-Bench 4.0 agentic coding, it hit 39.2% where its predecessor scored 0.0%. Knowledge-work scores: 1620 on GDPval-AA v2.1 (Haiku 4.5: 735; GPT-6 Luna: 1437) and 45.9% on Humanity's Last Exam without tools, 57.4% with tools (predecessor: 10.2% / 18.7%).
Early customer testing lines up with the lab numbers. Anthropic quotes Asana staff engineer Aaron Vinh saying his team saw over 30% lower latency on task completions and up to 2.5x faster inference per agent turn across their AI Teammates eval suite. Reuters, meanwhile, frames the timing as portfolio-building ahead of a planned IPO.
Safety and where it fits
Haiku 5.5 is the first Haiku with built-in restrictions for a narrow set of high-risk cybersecurity requests. Per the announcement, its safeguards permit a wider range of defensive tasks than Sonnet 5.5's while still blocking penetration testing and techniques more likely to be used by attackers. Biology safeguards match Sonnet 5, 5.5, and Opus 5. Anthropic says alignment evaluations show far fewer instances of misaligned behavior than Haiku 4.5.
Positioning matters here: Anthropic is explicit that Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding, while Haiku 5.5 is aimed at narrowly scoped tasks that were previously cost-prohibitive to run at scale — compaction, summarization, subagent work, live support, browser use. If your workload is thousands of small model calls a day, this is the model that changes the economics; if it's one hard problem, the bigger models still win.