← All insights

News analysisLens: United States3 min read

Claude Haiku 5.5 costs 90% less below 100,000 tokens. That line is now an architecture decision

Anthropic's Haiku 5.5 makes high-volume agent steps far cheaper, but only while a request stays under 100,000 tokens. Routing, compaction and caching now decide most of the bill.

Listen to this article · 5 min

AI-generated narration of the full article.

Measuring spoons of different sizes beside a pile of sugar, in La Madre duotone, beside the words Right size per step
Photo: Flat Lay Photos (StockSnap, CC0)

Ten cents. That is what Anthropic now charges for a million input tokens sent to Claude Haiku 5.5, released on October 7, as long as the request stays at or under 100,000 tokens. Cross that line and the same token costs five times as much.

Both numbers matter. The first changes which model an agent should call by default. The second turns the size of each request into something an architect has to design, not just observe.

What a million tokens costs, by model and request size
US$ per million tokensInputOutputCache read
Haiku 5.5, up to 100K tokens0.100.500.01
Haiku 5.5, above 100K tokens0.502.500.05
Haiku 4.51.005.000.10
Sonnet 5.52.0010.000.10
List prices published by Anthropic on October 7. Sonnet 5.5 cache reads were cut from $0.20 to $0.10 the same day.

The cheap call stops being a compromise

Haiku 4.5 cost $1 per million input tokens and $5 per million output. Haiku 5.5 costs a tenth of that below the line, and Anthropic describes it as about 75% cheaper to run overall, a blended estimate that mixes both price tiers. The model ID is claude-haiku-5-5, and it is available on Anthropic’s own platform, AWS, Google Cloud and Microsoft Azure. Databricks added it to Unity Gateway the same day.

Anthropic points it at exactly the steps that make up most of an agent’s traffic: classification, summarization, context compaction, database queries, subagents, live support and repetitive browser work. Its benchmark figures are Anthropic’s own, and they should be treated as a reason to test, not as a result.

We have argued that an agent should not ask a frontier model every small question. With Sonnet 5.5 at twenty times Haiku 5.5’s input price below the line, that argument stops being an optimization and becomes the default. The sensible starting point for a new agent is now the reverse of the usual one: every step begins on the small model, and a step is promoted to a larger model only when evaluation shows the small one failing it.

A price line inside the context window

The 100,000-token threshold is where design comes in. Above it, Haiku 5.5 is still half the price of Haiku 4.5, but each token costs five times what it did a moment earlier.

Agents cross that line in predictable ways. An orchestrator passes the whole conversation history to every subagent. A retrieval step returns twenty documents when three would do. A long-running task accumulates tool outputs nobody removes. Each habit was tolerable when every request cost about the same. Now each one has a price.

Two design patterns follow.

Compaction becomes a cost control. A cheap model summarizing state so that the next call stays under the line is exactly the job Anthropic advertises for Haiku 5.5. The risk is that a summary loses the one detail a later step needed, so the measure is task success after compaction, not tokens saved.

Stable prefixes become nearly free. Cache reads cost $0.01 per million tokens below the line. System prompts, tool definitions and policy text that repeat on every call should sit at the start of the request, unchanged, so they are read from cache. The same logic now applies to Sonnet 5.5, whose cache reads were halved.

Where the 90% does not show up

A price per token is not a cost per result. If the cheap model needs a retry on one call in five, or escalates a third of its cases to a larger model, the saving shrinks fast. That is why cost per accepted output, not per call, is the number to manage across a model portfolio.

Haiku 5.5 also ships with an effort setting from Low to Max. More effort usually means more output tokens, and on every row of the table above an output token costs five times an input token. Effort is a dial to set per step, not once per agent.

Finally, the path matters as much as the price. On Amazon Bedrock the model is offered through geographic inference profiles for the US, the EU, Australia and Japan, plus a global profile. A US enterprise with contractual commitments about where processing happens should choose the US profile deliberately rather than inherit the global default from a code sample.

The headline number is 90%. The useful number is a different one: the share of your agent’s calls that genuinely need anything bigger than the cheapest model you trust. Most teams have never measured it, and Anthropic has just made it worth measuring.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case