ProcureBench
LLM benchmark on procurement work
ProcureBench runs each model stack through the same three production agents on one school renovation tender, in English and in Latvian. Outputs are scored against a fixed answer key. Cost is the LLM usage of the whole cycle.
Composite score
Mean of the three activity scores, averaged over English and Latvian. Scale 0 to 100. Definitions are in the Method section.
Composite score by language
Gap is the English composite minus the Latvian composite. A negative gap means the stack scored higher in Latvian.
- English
- Latvian
Scores by activity
Each activity is scored 0 to 100 per language.
| Stack | Requirements extraction | Bid analysis | Market research | |||
|---|---|---|---|---|---|---|
| English | Latvian | English | Latvian | English | Latvian | |
|
Kimi K3
|
92.2 | 91.5 | 91.5 | 93.6 | 80.0 | 56.7 |
|
Claude tiered
|
90.2 | 81.1 | 95.4 | 95.0 | 70.0 | 70.0 |
|
GLM-5.3-Flash
|
80.3 | 92.1 | 95.2 | 91.2 | 76.7 | 54.4 |
|
Muse Spark 1.3
|
90.8 | 81.6 | 88.6 | 95.0 | 60.0 | 70.0 |
|
Gemini 3.8 Flash
|
91.4 | 82.6 | 100.0 | 90.9 | 60.0 | 60.0 |
|
DeepSeek V4.1 Flash
|
89.7 | 91.2 | 87.5 | 87.6 | 66.1 | 50.0 |
|
OpenAI tiered
|
92.2 | 92.4 | 82.7 | 83.6 | 53.3 | 53.3 |
Cost, time and reliability
Cost is the LLM usage of both languages and all activities, including retries. Wall-clock is the duration of the whole cycle. Reliability is completed activity records divided by attempts.
| Rank | Stack | Composite | English | Latvian | EN − LV | Total cost | Wall-clock | Reliability | Trials |
|---|---|---|---|---|---|---|---|---|---|
| 1 |
Kimi K3
OpenRouter, reasoning xhigh main / medium mid / minimal fast
|
84.3 | 87.9 | 80.6 | +7.3 | $46.52 | 5 h 57 min | 100% | 1 |
| 2 |
Claude tiered
Opus 4.8 main, Sonnet 4.6 mid, Haiku 4.5 fast
|
83.6 | 85.2 | 82.0 | +3.2 | $77.69 | 7 h 7 min | 100% | 1 |
| 3 |
GLM-5.3-Flash
OpenRouter, all tiers
|
81.6 | 84.1 | 79.2 | +4.9 | $4.15 | 13 h 38 min | 89% | 1 |
| 4 |
Muse Spark 1.3
OpenRouter, all tiers
|
81.0 | 79.8 | 82.2 | -2.4 | $39.80 | 2 h 11 min | 100% | 1 |
| 5 |
Gemini 3.8 Flash
all tiers
|
80.8 | 83.8 | 77.8 | +6.0 | $70.54 | 6 h 7 min | 100% | 1 |
| 6 |
DeepSeek V4.1 Flash
OpenRouter, all tiers
|
78.7 | 81.1 | 76.3 | +4.8 | $13.90 | 3 h 14 min | 100% | 1 |
| 7 |
OpenAI tiered
GPT-5.6 Sol main, GPT-5.6 Terra mid, GPT-5.4 mini fast
|
76.3 | 76.1 | 76.5 | -0.4 | $30.21 | 2 h 9 min | 100% | 1 |
Stacks, models and judges
| Stack | Main role | Mid role | Fast role | Judge | Run date |
|---|---|---|---|---|---|
|
Kimi K3
|
kimi-k3 | kimi-k3 | kimi-k3 | claude-fable-5-1 | 2026-09-07 |
|
Claude tiered
|
claude-opus-4-8 | claude-sonnet-4-6 | claude-haiku-4-5 | claude-fable-5-1 | 2026-09-05 |
|
GLM-5.3-Flash
|
glm-5.3-flash | glm-5.3-flash | glm-5.3-flash | claude-fable-5-1 | 2026-09-07 |
|
Muse Spark 1.3
|
muse-spark-1.3 | muse-spark-1.3 | muse-spark-1.3 | claude-fable-5-1 | 2026-09-05 |
|
Gemini 3.8 Flash
|
gemini-3.8-flash | gemini-3.8-flash | gemini-3.8-flash | claude-fable-5-1 | 2026-09-05 |
|
DeepSeek V4.1 Flash
|
deepseek/deepseek-v4.1-flash | deepseek/deepseek-v4.1-flash | deepseek/deepseek-v4.1-flash | gpt-6 | 2026-09-11 |
|
OpenAI tiered
|
gpt-5.6-sol | gpt-5.6-terra | gpt-5.4-mini | claude-fable-5-1 | 2026-09-06 |
Method
A stack is the set of models assigned to the three agent roles: main, mid and fast. Different steps of each agent call different roles. A stack may use one model for all three roles or a different model per role.
The corpus is one municipal school renovation tender with two lots, in parallel English and Latvian versions: regulations, a technical specification, a bill of quantities, a draft contract, a drawing and two bids. One bid is compliant. The other contains 22 planted defects; some appear only in the spreadsheet, some only on scanned certificate pages.
For each stack, language and trial, the benchmark runs requirements extraction on the RFP once, bid analysis once for each of the two bids, and market research on the tender's items once. It uses the same jobs, sandboxes and tools as production, with one exception: every catalogue item is approved automatically so market research can start without a human step.
Bid analysis is scored on planted-defect recall (50%), precision against false alarms on either bid (30%) and verdict correctness (20%); the trial score is the lower of the two bids. Requirements extraction is scored on reference recall (60%), precision (20%) and the lots and criteria structure (20%). Market research is 40% structural (run completed, share of items with at least one quote, share of quotes with a validated source) and 60% a rubric of facts the report must contain. With more than one trial, each activity score is the median across trials. The language composite is the mean of the three activity scores; the overall composite is the mean of the English and Latvian composites.
A separate model reads each output against the answer key in a worksheet; the scorer then applies the fixed weights above. The judge is never one of the models under test, but it may come from the same vendor. The judge for each stack is listed in the stacks table.
Cost is LLM usage only. First-party calls are priced at the provider list price on the run date, with cache reads, cache writes and thinking tokens weighted as the provider bills them. OpenRouter calls use the cost OpenRouter reported. Wall-clock is the time from driver start to finish for both languages, including setup and queue wait.
Reliability is completed activity records divided by the number of attempts; a retry that later succeeds counts as two attempts and one completion. A stack with one trial is a single sample.