One model, two price lists: how Haiku 5.5 prices agent context
Claude Haiku 5.5 costs a tenth of its predecessor under 100k prompt tokens and five times more past them. Where an agent budget lands depends on which side of that line your prompts live on.
Anthropic shipped Claude Haiku 5.5 on October 7 with the usual benchmark chart and one row that deserves more attention: a pricing table split in two. Under 100,000 prompt tokens the model costs a tenth of its predecessor, $0.10 per million input tokens and $0.50 per million output tokens. Past 100k, every one of those numbers is five times higher. One model, two price lists, and the boundary sits where agent work tends to start.
The launch headline says Haiku 5.5 runs about 75% cheaper than Haiku 4.5 on average. That average is honest and it hides the shape of the schedule. A note attached to the pricing table gives it away: around 90% of requests to the previous Haiku stayed under 100,000 prompt tokens. For that majority the sticker price holds and the discount is real. The remaining slice, the long-transcript requests where an agent drags a whole codebase into context, gets billed from the other column.
The multiplier applies to all four line items. Input goes $0.10 to $0.50, output $0.50 to $2.50, cache reads $0.01 to $0.05, cache writes $0.125 to $0.625. The table picks a column by prompt size, up to 100k or over, and a request bills at the rate its column sets. At long-column rates, 150,000 input tokens cost 7.5 cents a call; five calls an hour across a working day is three dollars of input alone. Small numbers that compound the moment a workflow stops being a demo.
There is a second catch, buried in the migration guide: the tokenizer changed too. The same input text produces roughly 30% more tokens on Haiku 5.5 than on Haiku 4.5. A prompt that sat at 80k on the old model can cross the 100k line after the recount, and the discount you budgeted quietly re-prices itself into the five-x column. Price per token dropped ninety percent, tokens per prompt rose thirty percent, and the boundary moved closer to your requests at the same time. Measure prompt lengths with the new model id before assuming which column you live in.
The boundary sits at 100k because that is where the cheap model's job ends and Sonnet's begins. Anthropic positions Haiku 5.5 for high-volume, cost-sensitive work: summaries, compaction, database queries, classification runs. Those live under the line. The long-column prices acknowledge the overlap: $2.50 per million output tokens is a quarter of Sonnet 5.5's $10.00, and $0.05 for cache reads is exactly half of Sonnet's rate after this week's cut. The discount narrows as context grows, on purpose.
The capability jump makes the structure defensible. On OSWorld 2.1, which scores how well agents drive a real computer through multi-step tasks, Haiku 5.5 scores 72.4% against 15.7% for Haiku 4.5. Anthropic's own table puts it ahead of GPT-6 Luna at 48.9% and behind Sonnet 5.5 at 83.9%. On Terminal-Bench 4.0 the previous Haiku scored 0.0%; this one completes 39.2% of tasks. A model that could not run the harness at all now finishes four in ten, at a price meant to keep you off the meter.
Two quieter moves shipped alongside. Sonnet 5.5 cache reads dropped from $0.20 to $0.10 per million tokens, which Anthropic estimates at roughly 20% off most agentic work since cache reads dominate token consumption. And Max and Team subscribers get monthly API credits this week: $100 for Max 5x, $200 for Max 20x, up to $500 pooled on Team, usable on any model. Both levers point at the same buyer, the builder running agent sessions every day, and each one moves a different line of that bill.
The switch itself is one model id, claude-haiku-5-5, live on AWS, Google Cloud, and Azure with no date suffix and no alias. The guide around it reads like a list of 400 errors waiting to happen: budget-style thinking configs, sampling parameters, and assistant prefills all hard-fail on the new model. That work comes before any cost math. Measure your prompt-length distribution on the new tokenizer for a week. If every call sits under 100k after the recount, the move is free money and Haiku 4.5 has no reason to stay in your config. If every task opens with one 200k-token sweep, you will pay Sonnet-adjacent rates on the expensive part, and the 90% headline was never about you. There is also an effort knob, a first for the Haiku line, so a per-call choice between cost and depth now exists where a model swap used to be the only lever.
The system card reports fewer misaligned behaviors than 4.5 across Anthropic's evals, with cybersecurity safeguards somewhere between the two generations: wider than Sonnet 5.5 on defensive work, still closed to offense. For teams routing untrusted input through small models, that page is worth reading before the price is.
Pricing used to be a row in a table you skimmed once. Now the row forks at a token count, and the fork is where agent budgets land. Haiku 5.5 earns its slot; the capability jump is real and the short-column rate is aggressive by any baseline. Know which column your p99 lives in before the first invoice counts it for you.
Comments ()