Anthropic launched Claude Haiku 5.5 on October 7, 2026, the smallest and fastest model in the Claude 5.5 generation. The focus is high-volume, cost- and latency-sensitive tasks: summaries, classification, extraction, subagents, and real-time support.
Pricing and performance
For prompts up to 100,000 tokens — about 90% of previous Haiku usage — the price drops to $0.10 per million input tokens and $0.50 output. That is roughly 90% less than Haiku 4.5 in that range, or on average 75% less accounting for the new tokenizer. Longer prompts cost $0.50 / $2.50.
In Anthropic's benchmarks, Haiku 5.5 shows clear jumps: OSWorld 2.1 (computer use) rises from 15.7% on Haiku 4.5 to 72.4%; Terminal-Bench 4.0 goes from 0% to 39.2%; Humanity’s Last Exam (with tools) from 18.7% to 57.4%. On knowledge work (GDPval-AA) it stays close to Sonnet 5.5 on some metrics and beats GPT-6 Luna on several items.
It is the first Haiku with adjustable effort control, allowing a balance between cost and intelligence. Anthropic also halved the cache-read price for Sonnet 5.5 and announced monthly API credits for Max and Team subscribers.
Why it matters
With the price cut, it becomes practical to use Haiku 5.5 as a subagent in pipelines that previously required larger models — for example, an Opus or Sonnet plans and Haiku executes search, summary, or browser-use tasks in parallel. Customers such as Asana, HubSpot, and Cognition reported latency and cost gains in early tests.
The model is available on the Claude Platform, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. 1-million-token context and up to 128K output (300K in batch beta).
Sources
Transparency: This content was created, edited or reviewed with the assistance of artificial intelligence. The information was cross-checked with public posts on X and sources available on the internet. Consult the original sources to verify the full context.
By GeekikiBot