Skip to main content
We switched Tendergate to Haiku 5.5: the reading bill fell 75%, and verdicts on the same bid got stricter
October 08, 2026

We switched Tendergate to Haiku 5.5: the reading bill fell 75%, and verdicts on the same bid got stricter

Between 10 September and 8 October we analysed the same public tender eight times: 53 checkable requirements, 22 award criteria, several versions of one real bid. Tender, buyer and bidder stay unnamed. In the first seven runs, Haiku 4.5 did the item-by-item reading, and that stage cost $14.39 to $16.04 every time. On 8 October, the day after Anthropic released Haiku 5.5, we switched Tendergate to it. The same stage then cost $3.96, 75% less — matching Anthropic's launch claim of “around 75% less to run” than Haiku 4.5.

−75%
reading-stage cost on the same bid, $16.04 to $3.96
17 min
reading-stage time, against 24–38 minutes on Haiku 4.5
14 / 53
requirements where the verdict changed on the identical bid

Our bid analysis uses three tiers of models. The fast one reads item by item: one small agent per requirement, criterion and lot, 76 on this tender.

1. The bill

Reading-stage cost per analysis, same tender
Runs 1–7 on Haiku 4.5 (10–24 September), run 8 on Haiku 5.5 (8 October).
Run 1 · 4.5 $14.39
Run 2 · 4.5 $14.67
Run 3 · 4.5 $16.01
Run 4 · 4.5 $15.22
Run 5 · 4.5 $15.43
Run 6 · 4.5 $15.06
Run 7 · 4.5 $16.04
Run 8 · 5.5 $3.96

The seven Haiku 4.5 runs landed within $1.65 of each other, so the drop is not noise. The whole analysis fell from $36.27 to about $18, and the main model, at $8.64, is now its largest cost.

2. More thinking per step

Smarter models are supposed to need fewer steps — Anthropic said so when it launched Opus 4.5. Requirements took as many steps — model calls — as before; award criteria took about half as many. Every step thought far more.

11.1
calls per requirement on Haiku 5.5, against 9.8–17.8 on Haiku 4.5
6.0
calls per award criterion, against 11.0–18.1
13×
more thinking per call on the same bid

Prompts grew too, to 80,000 tokens against 52,000, partly because Haiku 5.5 counts the same text as about 30% more tokens. That matters because of a price line: below 100,000 tokens per request, Haiku 5.5 costs a tenth of Haiku 4.5; above it, the whole request costs five times as much. 187 of our 724 calls crossed the line: a quarter of the calls, 68% of the stage's bill.

3. Stricter on the same bid

Runs 7 and 8 read the identical bid version:

Requirement verdictHaiku 4.5, 24 SeptemberHaiku 5.5, 8 October
Compliant4935
Partially compliant06
Could not be verified18
Not applicable34

All 14 changes moved away from "compliant". Some are better catches, such as a submission requirement the older run passed with no upload receipt in the bid. Some are too strict, such as two exclusion grounds the buyer checks itself in public registers, left unverified. On award criteria it went the other way: three of four changes upgraded a partial score to compliant.

Not all of this is the model. Between the runs we added an instruction that only the bidder's own entries count as proof, moved the model that re-checks findings from Sonnet 4.6 to Sonnet 5.5, and changed how documents are turned into text. Most of the cost drop is the model's price; the verdict changes belong to the whole pipeline, and one run cannot separate them.

4. What we take from it

Keep each agent's prompt under 100,000 tokens, and re-run your own documents before trusting a model swap, as Anthropic's Haiku 5.5 prompting guide also advises.

Every verdict in Tendergate carries the quote and document location it rests on, so any verdict that changes after a model swap can be checked directly against its source.

How we counted: eight production analyses; main model Opus 4.8 in runs 1–6 and Opus 5.5 in runs 7–8. Costs use Anthropic's published prices, including the higher rate above 100,000 tokens.

See how AI can help with your procurements.
Try your first AI analysis for free.
Register for free

Sources

  1. Introducing Claude Haiku 5.5 (Anthropic, 7 October 2026). Release date and the "around 75% less to run" claim.
  2. Claude API pricing (Anthropic). Haiku 5.5 and Haiku 4.5 prices, including Haiku 5.5's higher rate above 100,000 tokens.
  3. Migrating to Claude Haiku 5.5 (Anthropic). The same text counting as roughly 30% more tokens on Haiku 5.5.
  4. Prompting Claude Haiku 5.5 (Anthropic). Advice to compare effort settings on your own evaluations.
  5. Introducing Claude Haiku 4.5 (Anthropic, 15 October 2025). Haiku 4.5's $1 and $5 per million token pricing.
  6. Introducing Claude Opus 4.5 (Anthropic, 24 November 2025). The claim that smarter models solve problems in fewer steps.
  7. Introducing Claude Opus 5.5 (Anthropic, 22 September 2026). The main model in runs 7 and 8.
  8. Introducing Claude Sonnet 5.5 (Anthropic, 28 September 2026). The model that re-checked findings in run 8.
  9. Introducing Sonnet 4.6 (Anthropic, 17 February 2026). The model that re-checked findings in runs 1 to 7.
How this post was written

We build AI for procurement, so it would be a little odd not to use it here. This post was written together: AI for the tireless reading and first drafts, people for the judgement, the corrections and the final yes.

Back to blog