⏱️ Reading time: 9 min

Anthropic cut the price of its smallest model by 75% on average and, at the same time, made it the fastest the company offers as of October 2026. Claude Haiku 5.5 launched on October 7 as a direct replacement for Haiku 4.5, with an added advantage: for the first time, a model in the Haiku line lets users adjust how much reasoning effort it applies to each task.

📑 En este artículo
  1. TL;DR
  2. What Claude Haiku 5.5 Is
  3. What Happened
  4. Context and Background
  5. Technical Details: How Effort Tuning Works
  6. Impact and Analysis
  7. What’s Next
  8. Frequently Asked Questions
    1. What is Claude Haiku 5.5 and how does it differ from Haiku 4.5?
    2. How does Haiku 5.5’s effort tuning work?
    3. Does Haiku 5.5 replace Sonnet 5.5 or Opus 5.5?
    4. How much cheaper is the lightweight model in the Claude family compared to its previous version?
    5. What’s changing in Claude Sonnet 5.5’s pricing?
  9. References

The launch comes alongside two other pricing changes: cache reads for Claude Sonnet 5.5 now cost half as much, and Claude Max and Team subscribers get a new monthly credit to build agents on the Claude Platform.

TL;DR

  • Anthropic launched Claude Haiku 5.5 on October 7, 2026, with an average cost 75% lower than Haiku 4.5.
  • The new model adds configurable effort tuning (Low to Max), previously exclusive to Opus and Sonnet.
  • Sonnet 5.5’s cache reads dropped by half: agentic work now costs around 20% less.
  • Haiku 5.5 scores 72.4% on OSWorld 2.1 versus 15.7% for Haiku 4.5, according to the computer-use benchmark.
  • Claude Max and Team now include a monthly API credit for building new agents on the platform.

What Claude Haiku 5.5 Is

Claude Haiku 5.5 is the smallest and most affordable model in Anthropic’s Claude 5.5 family, built for high-volume tasks like summarization, classification, and database queries. It also works as a coding subagent alongside Opus 5.5 and Sonnet 5.5.

According to Anthropic’s official announcement, the model is optimized for speed-sensitive work, such as live customer support and browser use, and is the fastest model the company has released to date.

Anthropic introduced Haiku 5.5 as the fastest and cheapest model in its lineup as of October 7, 2026. Foto de Brecht Corbeel en Unsplash

What Happened

On October 7, 2026, Anthropic released Claude Haiku 5.5 as the successor to Haiku 4.5. The company describes it as the cheapest, fastest, and most capable small model we’ve ever released on its launch page. The model is already available on the Claude Platform, with an average cost 75% lower than its predecessor.

Alongside the model, Anthropic announced two additional pricing changes. First, cache reads for Claude Sonnet 5.5 dropped by half, making that model around 20% cheaper for most agentic work. Second, Claude Max and Team subscribers now get a new monthly API credit, meant to help them build their own agents and applications on the platform.

Haiku 5.5 also introduces a feature once exclusive to the larger models: configurable effort tuning. Users can choose between Low, Med, High, Xhigh, and Max levels depending on whether they prioritize cost or accuracy for a given response, the same approach already used by Opus 5.5 and Sonnet 5.5.

Context and Background

The Haiku line has been the smallest in the Claude family since its first version. While Opus targets deep reasoning and Sonnet aims for a middle ground between cost and capability, Haiku has always been positioned as the volume option: thousands of calls per minute at low prices. Haiku 4.5, its immediate predecessor, already held that role in high-traffic products, but lagged far behind the larger models on tasks requiring multiple reasoning steps or tool use.

That capability gap is what changes with Haiku 5.5. On Terminal-Bench 4.0, which measures agentic coding in a real terminal, Haiku 4.5 scored 0.0%. Haiku 5.5 reaches 39.2%, still far from Sonnet 5.5’s 70.6%, but now in a range where the small model can complete entire tasks, not just code fragments.

Market context matters too. Clients like Asana, HubSpot, AlphaSense, and Box tested Haiku 5.5 before launch, and their results are cited directly in the announcement. That kind of early enterprise validation has become standard practice in Anthropic’s recent releases, as the company looks to show measurable impact before opening a model to the general public.

Haiku 5.5’s effort tuning reuses the same Low-Max scale already found in Opus 5.5 and Sonnet 5.5. Foto de Planet Volumes en Unsplash

Technical Details: How Effort Tuning Works

The core technical point of this launch isn’t the model itself, but the effort control now shared across the entire Claude 5.5 family. In practice, the effort parameter decides how many internal reasoning tokens the model generates before responding: at Low, Haiku 5.5 prioritizes speed and spends few thinking tokens; at Max, the model invests more compute per response and gets closer to the accuracy of a larger model, at the cost of latency and price. It’s also the first model in the Haiku line with this option: neither Haiku 4.5 nor any earlier Haiku allowed users to choose how much to reason before responding.

That mechanism explains why Haiku 5.5 scores 45.9% on Humanity’s Last Exam without tools and, on the same benchmark, gets closer to Sonnet 5.5 (56.9%) when given more effort: it’s not that the model knows more at Max, it dedicates more reasoning steps to checking its own answer before delivering it. That cost isn’t free: according to the accuracy-versus-cost charts Anthropic published, moving from Low to Xhigh on OSWorld 2.1 multiplies the cost per task several times over for a few percentage points of accuracy, so it’s worth reserving the higher levels for tasks where errors are expensive.

GDPval-AA v2.1, designed by Artificial Analysis, evaluates agents on real professional work across 44 different occupations, and is another benchmark that shows the gap between models:

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5 (reference)
GDPval-AA v2.1 (Elo)162073514371840
AA-Briefcase v1.1 (Elo)157861413361824
OSWorld 2.1 (offline, computer use)72.4%15.7%48.9%83.9%
Humanity’s Last Exam (no tools)45.9%10.2%, 56.9%
Terminal-Bench 4.0 (agentic coding)39.2%0.0%16.4%70.6%

The numbers come from Anthropic’s own Haiku 5.5 system card. GPT-6 Luna appears as an external reference; Sonnet 5.5 is included not because it competes at the same price point, but to show how much capability gap remains between the small model and the mid-size model from the same company.

flowchart LR
A["Input task"] --> B{"Effort level"}
B -->|"Low"| C["Fast response, minimal cost"]
B -->|"Med"| D["Balance of cost and accuracy"]
B -->|"High"| E["More internal reasoning tokens"]
B -->|"Xhigh or Max"| F["Maximum accuracy, higher cost per task"]

The diagram sums up the idea: effort doesn’t change the model, it changes how much it thinks before responding.

Impact and Analysis

The most immediate impact is economic: if a company was already running millions of daily calls with Haiku 4.5 for classification or summarization, the same volume now costs a fraction of what it did before. AlphaSense reported that its Ask in Document product, which processes about 8 million queries per week, went from a score of 0.76 to 0.84 with Haiku 5.5 in a 400-query test, an improvement the company called statistically significant.

Box, for its part, measured a score 11 points higher than Haiku 4.5 with half the latency, and applied it to analytical work at scale, such as cost reports and financial summaries. Asana reported more than a 30% reduction in latency and up to 2.5 times faster inference per agent turn in its evaluation suite for AI Teammates. HubSpot, which evaluates small models on simulated CRM tasks, said Haiku 5.5 achieved the best score it has seen on that test so far: 92.8% averaged across three runs.

💭 Key takeaway: Anthropic didn’t improve Haiku by selling more intelligence at the same price: it sold roughly the same intelligence at a fraction of the price, and saved the capability gains for those willing to pay for more effort.

Taken together, these figures point to a shift in strategy rather than a one-off leap: Anthropic is pushing Claude Haiku 5.5 into tasks that used to require a mid-size model, and it’s doing so by making the small model cheaper instead of making the large one more expensive. That puts pressure on competitors to respond with their own budget models, in a category where margins are decided by cents per million tokens, not frontier benchmarks.

What’s Next

Anthropic gave no date for the next update to the Haiku line, but the pattern of recent releases suggests short cycles: Haiku 4.5 arrived before it, and now Haiku 5.5 extends effort tuning across the entire catalog. The next generation of Sonnet and Opus will likely inherit similar efficiency gains before Haiku gets updated again.

For developers, the practical change worth watching is the monthly API credit for Claude Max and Team subscribers: if that credit is enough to run real agent projects without paying separately, it changes the calculation of which plan makes sense for small teams that currently pay for API usage apart from their subscription. Anthropic didn’t publish the exact credit amount in the announcement, so that detail will need to be confirmed once billing reflects it.

Try it yourself: go to the Anthropic console and run the same classification task with Haiku 4.5 and Haiku 5.5 to compare cost and latency with your own data.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What is Claude Haiku 5.5 and how does it differ from Haiku 4.5?

Claude Haiku 5.5 is the latest version of Anthropic’s small model, with an average cost 75% lower and speed and accuracy improvements over Haiku 4.5, according to benchmarks the company published on October 7, 2026.

How does Haiku 5.5’s effort tuning work?

Effort tuning lets users choose between Low, Med, High, Xhigh, and Max levels. The higher the level, the more internal reasoning tokens the model uses before responding, which raises accuracy but also cost and latency per task.

Does Haiku 5.5 replace Sonnet 5.5 or Opus 5.5?

No. Anthropic positions it as a subagent alongside Opus 5.5 and Sonnet 5.5 for high-volume tasks, not as a substitute for work that requires the deeper reasoning of the larger models.

How much cheaper is the lightweight model in the Claude family compared to its previous version?

Anthropic reports an average cost 75% lower than Haiku 4.5, though the final savings depend on the effort level chosen and the token volume of each task.

What’s changing in Claude Sonnet 5.5’s pricing?

Sonnet 5.5’s cache reads dropped by half, making that model around 20% cheaper for most agentic work, according to the same announcement.

References

📱 Enjoying this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day. @programacion

Featured image: Foto de Igor Omilaev en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.