Claude Haiku 5.5 is designed for work that needs to be fast, frequent, and economical—without being limited to simple chat. Anthropic launched the model on October 7, 2026, positioning it for high-volume tasks such as summaries, classification, database queries, live customer support, browser use, and sub-agent work.
The practical takeaway is to test it on well-scoped workloads where latency and cost matter. Anthropic’s benchmark and customer results are useful starting points, but they do not predict how the model will perform on your specific data or workflow.
Claude Haiku 5.5 at a glance
- Model ID:
claude-haiku-5-5 - Positioning: Anthropic’s fastest and cheapest small model to date, aimed at high-volume, cost-sensitive work.
- Context and output: a one-million-token context window and up to 128,000 output tokens.
- Effort setting: adaptive thinking with an effort control, allowing a choice between lower cost and more capability.
- Availability: Anthropic says it is available now across its platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure.
- Computer and browser use: beta support is being added to the Claude Python and TypeScript SDKs.
API pricing and cache changes
Anthropic lists the following per-million-token rates for Haiku 5.5:
| Request size | Input | Output | Cache reads | 5-minute cache writes |
|---|---|---|---|---|
| Up to 100,000 tokens | $0.10 | $0.50 | $0.01 | $0.125 |
| Over 100,000 tokens | $0.50 | $2.50 | $0.05 | $0.625 |
The company says Haiku 5.5 is priced 90% below Haiku 4.5 for requests up to 100,000 tokens and 50% lower above that size. Anthropic’s estimated average savings also account for changes in token counts from an updated tokenizer, so teams should compare actual cost per completed task rather than infer savings from a single rate.
Anthropic separately cut Sonnet 5.5 cache-read pricing by half, from $0.20 to $0.10 per million tokens. It estimates this makes Sonnet 5.5 around 20% cheaper on most agentic work. That reduction applies to cache reads, not all token types.
The launch also coincides with monthly API credits for eligible Max and Team subscribers. Anthropic’s current help page lists $100 for Max 5x and $200 for Max 20x; Team credits are $20 per Standard seat and $100 per Premium seat, pooled up to $500 per month. These credits fund the Claude Platform API and agent tools; they do not add to a subscriber’s interactive Claude, Claude Code, or Cowork usage limits. Check eligibility, account linking, and current terms before relying on them.
Where the smaller model may fit
Anthropic recommends Haiku 5.5 for repetitive, narrowly scoped work and as a sub-agent alongside larger models. This can be a useful pattern when a main model handles planning or synthesis while a smaller model performs bounded lookups, extracts a field, classifies an item, or condenses material.
The model can also support latency-sensitive tasks such as live support and browser workflows. Anthropic says beta computer-use and browser-use capabilities are being added to its Python and TypeScript SDKs. Because these are actions in a tool environment, start with a limited sandbox, restrict permissions, log the actions, and require approval for consequential changes.
Do not choose from benchmarks alone
Anthropic publishes results across knowledge-work, computer-use, reasoning, coding, and visual-reasoning evaluations. The company also shares early customer examples. These results can help identify candidate tasks, but test conditions and workloads differ. A benchmark result is not a service-level promise or evidence that a model is suitable for a specific production decision.
Haiku 5.5’s safeguards also vary by risk area. Anthropic says its cybersecurity safeguards are more restrictive than Haiku 4.5’s but less restrictive than those for Sonnet 5.5; the model still blocks penetration testing and other techniques it considers more likely to be misused. Review the current system card and provider policy before planning security research with the model.
Migration details can break a model-ID-only upgrade
Anthropic’s migration guide lists several changes from Haiku 4.5. The same text now produces about 30% more tokens because Haiku 5.5 uses a newer tokenizer, so old max_tokens limits and cost estimates need to be revisited. Adaptive thinking is the new approach; requests using the former enabled-thinking configuration with a budget_tokens value need updating, and effort controls can help manage how much the model reasons.
The guide also says to omit temperature, top_p, and top_k, replace assistant-prefill endings, and select response content blocks by their type rather than assuming the answer is the first block. On the Claude API and Google Cloud, computer use also moves to the computer_toolset_20260801 toolset; browser use uses a separate toolset. Stored thinking blocks must be replayed through the account that produced them, and earlier conversation turns must remain unchanged when those blocks are sent back. Treat the official migration checklist as part of the upgrade, not optional polish.
A practical evaluation plan
- Select representative, low-risk tasks that currently run at meaningful volume.
- Define success before testing: acceptable accuracy, response time, escalation rate, and maximum cost per successful result.
- Compare Haiku 5.5 with your current model using the same inputs and tool permissions.
- Include retries, cache behavior, longer prompts, and human corrections in cost calculations.
- Pilot behind a reversible routing rule, keep logs, and retain a human escalation path.
For teams integrating model APIs or agent workflows, AI consulting and custom AI development can help build a representative evaluation and deployment plan.
Official sources
Frequently asked questions
What is Claude Haiku 5.5?
Claude Haiku 5.5 is Anthropic’s small, fast model for high-volume and cost-sensitive tasks such as summarization, classification, database queries, live support, browser use, and agent sub-tasks.
What is the Claude Haiku 5.5 API model ID?
Anthropic lists the model ID as claude-haiku-5-5. Confirm the current model documentation and your provider’s availability before changing an integration.
How much does Claude Haiku 5.5 cost?
Anthropic lists per-million-token prices of $0.10 input and $0.50 output for prompts up to 100,000 tokens; for prompts over 100,000 tokens the listed rates are $0.50 input and $2.50 output. Cache rates differ by prompt length. Check Anthropic’s current pricing before estimating a workload.
Can Haiku 5.5 use a computer or browser?
Anthropic says its Claude Python and TypeScript SDKs are adding computer-use and browser-use support in beta. Confirm current SDK requirements, tool permissions, and supported platforms before deployment.
Should I replace another Claude model with Haiku 5.5?
Not automatically. Use representative tasks to compare accepted results, latency, retries, token usage, and total cost. Anthropic says Haiku 5.5 is best suited to narrower high-volume tasks, while Sonnet and Opus remain better choices for some complex agentic coding work.
