Anthropic Launches Haiku 5.5 With Major Price Cuts and Performance Gains

Third 5.5 model release in a month brings 90% price reduction for small requests and strong agentic benchmark scores

By LineZotpaper
Published
Read Time2 min
Anthropic has released Claude Haiku 5.5, the first update to its smallest and most affordable model in nearly a year, slashing per-token prices by up to 90% for smaller requests while claiming dramatic improvements across agentic coding, computer use and reasoning benchmarks compared with its predecessor.

The new model, described by Anthropic as its fastest and most efficient, is priced at $0.10 per million input tokens and $0.50 per million output tokens for requests under 100,000 tokens, down from $1 and $5 respectively for Haiku 4.5. For larger requests, the cost is $0.50 input and $2.50 output, though Anthropic notes that roughly 90% of Haiku 4.5 requests fell into the lower-priced tier, putting average savings at about 75% when accounting for a mix of request sizes and an updated tokenizer.

Haiku 5.5 is the first Haiku model to offer effort controls, with a default setting of medium, giving developers more control over token usage per task. The model is intended for high-volume tasks such as summarisation, classification and routing, but Anthropic also highlights its suitability for agentic workloads where speed matters, including live customer support, browser use, compaction and database queries.

Benchmarks published by Anthropic show substantial improvements over Haiku 4.5. On the OSWorld 2.1 offline subset, which tests computer use, Haiku 5.5 scored 72.4%, up from 15.7% for its predecessor and ahead of OpenAI's GPT-6 Luna at 48.9%. On agentic coding benchmarks, Terminal-Bench 4.0 returned 39.2% (up from 0.0%) and FrontierCode 1.1 gave 46.4%. On visual reasoning (Chartography, no tools), the score was 46.4% versus 6.4% for Haiku 4.5. On the Knowledge work GDPval-AA v2.1 benchmark, Haiku 5.5 scored 1620, compared with 735 for Haiku 4.5 and 1437 for GPT-6 Luna.

The launch of Haiku 5.5 follows Sonnet 5.5 and another 5.5 series model released in the past month. Anthropic has not yet released Fable 5.5, which is expected to undergo a longer review process.

§

Analysis

Why This Matters

  • The price reduction makes Haiku 5.5 far more accessible for high-volume AI tasks, potentially reshaping the economics of deploying small models in production environments.
  • The large benchmark improvements, particularly in agentic coding and computer use, could broaden the role of small models beyond simple classification and routing to more complex, autonomous tasks.
  • Anthropic is competing aggressively on both capability and cost against OpenAI's GPT-6 Luna, setting a new pricing benchmark for the small-model tier.

Background

Anthropic's Claude model family spans three main sizes: Opus (largest), Sonnet (medium) and Haiku (smallest, fastest). The Haiku line has been used for high-throughput, latency-sensitive tasks. The 5.5 series marks a new generation of models, with Sonnet 5.5 arriving earlier and Haiku 5.5 being the third model in the series. Fable 5.5, the presumed next member, is still in development. The pricing shift is notable: previous Haiku models had a flat per-token rate, whereas Haiku 5.5 introduces a two-tier structure favouring smaller requests, and the effort control feature is new to the Haiku line.

Key Perspectives

Anthropic: Positions Haiku 5.5 as the fastest and most affordable model in its lineup, targeting developers who need high throughput for tasks like live customer support and browser automation, and emphasising the model's agentic capabilities. Competitors (e.g., OpenAI): GPT-6 Luna trails Haiku 5.5 on several published benchmarks, and the new pricing may force competitors to adjust their small-model pricing tiers or accelerate capability improvements to remain competitive. Developers and enterprises: Benefit from lower costs and improved performance, particularly for tasks that previously required larger models. However, switching from existing Haiku 4.5 workflows may require retesting and adjusting prompt strategies to take advantage of effort controls and the new tokenizer.

What to Watch

  • The share of Haiku 5.5 requests that fall under the 100,000-token threshold, which determines whether the lowest pricing tier applies.
  • Adoption rates and developer feedback on the effort control feature and its impact on output quality.
  • Timing and performance of Fable 5.5, and whether OpenAI responds with price cuts or a new small model to match Haiku 5.5's pricing and benchmark results.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.