The Ten-Cent Workhorse: Why Haiku 5.5 Rewrites the Economics of Agent Fleets
AI Tools & Automation

The Ten-Cent Workhorse: Why Haiku 5.5 Rewrites the Economics of Agent Fleets

The loudest AI launches still chase the biggest score on a leaderboard. The quieter one that landed on October 7, 2026, is the one that will show up on invoices. Anthropic released Claude Haiku 5.5 as its fastest small model yet and priced the common case so far below Haiku 4.5 that a lot of product teams will have to redraw the map of which model does which job. For prompts up to 100,000 tokens, input falls to $0.10 per million tokens and output to $0.50. That is a 90 percent cut versus the previous small model, and it is the band Anthropic says covered about 90 percent of Haiku 4.5 requests. The average running-cost claim is closer to 75 percent cheaper once a new tokenizer and longer prompts are included. That gap is not a coupon. It is a change in what is affordable to run all day.

TechMash has spent recent weeks on local silicon and open-weight efficiency, from the Surface Laptop Ultra shift toward desk-side agents to the workhorse gap in open-weight models. Haiku 5.5 is the cloud counterpart of that same argument. The frontier model is no longer the only place where capability moves. The cheap tier just stopped being a toy.

What actually changed in the price sheet

Anthropic's own announcement is the source that matters here, not a vendor slide. On the Claude Platform, Haiku 5.5 bills $0.10 per million input tokens and $0.50 per million output tokens when the prompt stays at or under 100,000 tokens. Cross that line and the rates step up to $0.50 input and $2.50 output. Cache reads are $0.01 or $0.05 per million depending on the same threshold. Cache writes land at $0.125 or $0.625. Haiku 4.5, by comparison, was a flat $1 input and $5 output, with cache reads at $0.10 and cache writes at $1.25. The short-prompt discount is therefore 90 percent on both input and output. The long-prompt discount is 50 percent. Anthropic's footnote is explicit that the headline 75 percent average also absorbs a tokenizer update, shared in spirit with Sonnet 5.5 and Opus 5.5, that uses slightly more tokens to finish the same piece of work. Anyone modeling a migration has to count tokens, not just list prices.

The company paired the launch with a second cut that does not get as many headlines. Cache reads on Claude Sonnet 5.5 drop from $0.20 to $0.10 per million tokens. Because cached context is a large share of what an agent actually consumes, Anthropic estimates that most agentic Sonnet workloads get about 20 percent cheaper. Haiku is the new subagent. Sonnet is the slightly less expensive lead. Together they change the shape of a multi-model stack more than either announcement does alone.

Rows of servers stand in for the always-on inference that a ten-cent small model suddenly makes routine.
Rows of servers stand in for the always-on inference that a ten-cent small model suddenly makes routine.

The scoreboard is no longer embarrassing

Price without competence is just a cheaper failure. The benchmark table Anthropic published with the launch is the reason this release is not a simple discount. On GDPval-AA v2.1, a knowledge-work suite, Haiku 5.5 scores 1620 against 735 for Haiku 4.5 and 1437 for GPT-6 Luna, with Sonnet 5.5 still ahead at 1840. On AA-Briefcase v1.1 the pattern repeats: 1578 versus 614 for the old Haiku and 1336 for Luna. Computer use is the sharper jump. On the offline subset of OSWorld 2.1, Haiku 5.5 reaches 72.4 percent, against 15.7 percent for Haiku 4.5 and 48.9 percent for GPT-6 Luna. Sonnet 5.5 remains the reference at 83.9 percent. Terminal-Bench 4.0, a command-line agent test, moves from 0.0 percent on Haiku 4.5 to 39.2 percent, ahead of Luna at 16.4 percent and well behind Sonnet at 70.6 percent. FrontierCode 1.1 main sits at 46.4 percent, a hair above Luna's 42.4 percent.

Those numbers should be read as a routing signal, not a victory lap. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding of the kind Terminal-Bench measures. Haiku 5.5 is aimed at the narrow, repeated jobs that used to be too expensive to attach to every session: compaction, summarization, classification, database lookups, and subagent calls. It is also the first Haiku-class model with an adjustable effort setting, so a team can spend more thinking on a hard turn and less on a routine one without swapping model IDs. That knob matters more at ten cents than it did at a dollar, because the floor is now low enough that extra effort is a choice rather than a budget event.

Circuit-level density is a useful picture of what changed: more useful work per token, not a bigger brand name.
Circuit-level density is a useful picture of what changed: more useful work per token, not a bigger brand name.

Why agent fleets feel this first

A single chat reply hides the economics. An agent fleet does not. A support product that classifies intent, pulls a policy snippet, drafts a reply, and then asks a larger model to approve the edge cases can multiply small-model calls by five or ten before a human ever sees the ticket. At Haiku 4.5 rates, teams rationed those calls. At Haiku 5.5 rates, the same pattern starts to look like infrastructure. Anthropic positions the model for live customer support and browser use precisely because it is the fastest Claude at standard speed, with the caveat that Opus in Fast Mode can still outrun it. Early customer notes in the launch post point the same direction. Asana reported more than a 30 percent latency cut on task completion and up to 2.5 times faster inference per agent turn in its AI Teammates evals. HubSpot said a CRM suite averaged 92.8 percent across three runs, with Haiku 5.5 fastest on an audit task that hunts stale but ambiguous records. AlphaSense, running about 8 million document-question calls a week, saw a statistically significant lift on a 400-query set. Box reported an 11-point gain over Haiku 4.5 at about half the latency. Cognition said Haiku 5.5 as a sidekick in Devin Fusion held a FrontierCode score of 66.2 while cutting cost and latency, with Opus 5.5 still in the lead role.

None of those quotes is an independent audit. They are still useful as a map of where buyers expect to put the model: high volume, short context, a bigger model nearby. That is the same architecture we have been tracking in consumer agents. Meta's Muse, covered in the quiet dossier piece on what an agent remembers, is a product that only works if background tasks are cheap enough to run without a user watching. A ten-cent classifier and summarizer is how those background loops stay solvent. The same logic applies to retrieval. EmbeddingGemma 2's pocket index lowers the cost of finding the right chunk. Haiku 5.5 lowers the cost of reading it and deciding what to do next. Pair them and a product can search, rank, and draft without calling a frontier model on every keystroke.

Agent work is a team sport: a small model routes and drafts while a larger one handles the turn that actually needs judgment.
Agent work is a team sport: a small model routes and drafts while a larger one handles the turn that actually needs judgment.

The credits, the clouds, and the migration trap

Availability is broad on day one. Free, Pro, Max, Team, and Enterprise users can select Haiku 5.5 on Claude.ai across web, iOS, and Android. Developers get it on the Claude Platform and through Amazon Web Services, Google Cloud, and Microsoft Foundry, plus Claude Code. The model ID is claude-haiku-5-5, with a 1 million token context window and up to 128,000 tokens of output. That context length is a gift and a trap. The gift is room for a long document. The trap is the price step at 100,000 tokens. A careless agent that stuffs an entire repository into every call will pay the higher band and erase much of the discount. Prompt caching is the intended escape hatch, and the new cache-read prices make that escape cheaper than it was on Haiku 4.5.

Anthropic is also handing Max and Team subscribers a monthly API credit meant for building, not for production burn. Max 5x accounts get $100, Max 20x accounts get $200, and Team plans get up to $500 pooled across users. The credits work on any Claude model. That is a distribution tactic as much as a discount: it lowers the cost of the first agent prototype so a team discovers Haiku 5.5 before a procurement cycle does. SDKs for Python and TypeScript are gaining beta support for computer use and browser use in the same window, which is the interface Haiku is being sold into. A builder who treats the credit as free production capacity will be surprised at month two. A builder who uses it to measure tokens per successful task will know whether the 100,000-token cliff is real for their workload.

There is a safety boundary worth stating plainly. Haiku 5.5's cybersecurity safeguards are tighter than Haiku 4.5's and looser than Sonnet 5.5's. Defensive work is allowed more broadly than on the midsize model. Penetration testing and similar attacker-leaning techniques stay blocked unless an organization is in the expanded Cyber Verification Program. Biology safeguards match Sonnet 5, Sonnet 5.5, and Opus 5: research questions pass, requests judged likely to cause harm do not, with a separate Life Sciences Verification Program for wider lab use. A cheap model that is also more capable is a dual-use story, and Anthropic is not pretending otherwise. Teams automating security review should read the system card before they point Haiku at production logs.

Global inference is a network problem: price cuts only matter if routing, caching, and safeguards travel with the call.
Global inference is a network problem: price cuts only matter if routing, caching, and safeguards travel with the call.

What to watch in the next quarter

Three tests will decide whether the ten-cent workhorse is a real platform shift or a launch-week spreadsheet. The first is mix. If most Haiku traffic stays under 100,000 tokens, the 90 percent cut survives contact with customers. If agents bloat their prompts, the effective discount collapses toward 50 percent and the tokenizer penalty bites harder. The second is routing quality. A fleet that sends hard Terminal-Bench-style coding to Haiku will look cheap and fail. A fleet that keeps Sonnet or Opus on the lead and Haiku on compaction, retrieval, and browser substeps should see the latency notes Asana and Box described without giving up the hard cases. The third is the competitive echo. Haiku 5.5's short-prompt rates match the band associated with GPT-6 Luna. If OpenAI, Google, or Mistral answer with another small-model cut, the floor moves again and the teams that instrumented cost per successful task will be the ones who can switch without a rewrite.

Local hardware does not disappear in this story. The DGX Spark desk-side superchip and the orbiting TPU experiments we covered in Google's first kilowatt-class satellite test are attempts to stop renting every inference call. Haiku 5.5 is the opposite bet made cheaper: keep renting, but rent a model that is finally good enough for the boring 80 percent. Most companies will do both. Sensitive context stays on a desk or a private cloud. Repetitive classification and support drafting go to the cheapest API that still clears an eval. The winner is not the lab with the single highest score. It is the stack that can name, for each step, why that model is the one being paid.

The bill is the product now

Haiku 5.5 will not replace Sonnet on the hardest coding sessions, and Anthropic does not claim it will. What it does is remove the excuse that small models are too weak or too expensive to sit inside every agent loop. A 1620 on GDPval-AA, a 72.4 percent OSWorld offline score, and a ten-cent input price on the common prompt length is a combination product teams can budget against this week. The official details live on Anthropic's Haiku 5.5 announcement and the Claude Platform Haiku 5.5 model overview developers use to migrate. Watch the 100,000-token step, the Sonnet cache-read cut, and whether early latency claims survive a month of production traffic. If they do, the next argument in AI will not be which model is smartest. It will be which model is cheap enough to run when nobody is watching.

Found this helpful? Share it!

Comments

0
No comments yet. Be the first!