Skip to main content

ProcureBench

LLM benchmark on procurement work

ProcureBench runs each model stack through the same three production agents on one school renovation tender, in English and in Latvian. Outputs are scored against a fixed answer key. Cost is the LLM usage of the whole cycle.

Updated September 11, 2026. Corpus version 2026-09-1, scoring version 1.

Stacks with one trial are single-run observations: their scores, ranks, gaps, cost and duration carry no variance estimate.

Composite score

Mean of the three activity scores, averaged over English and Latvian. Scale 0 to 100. Definitions are in the Method section.

84.3
83.6
81.6
81.0
80.8
78.7
76.3
Kimi K3
Claude tiered
GLM-5.3-Flash
Muse Spark 1.3
Gemini 3.8 Flash
DeepSeek V4.1 Flash
OpenAI tiered

Composite score by language

Gap is the English composite minus the Latvian composite. A negative gap means the stack scored higher in Latvian.

  • English
  • Latvian
English 87.9
Latvian 80.6
English 85.2
Latvian 82.0
English 84.1
Latvian 79.2
English 79.8
Latvian 82.2
English 83.8
Latvian 77.8
English 81.1
Latvian 76.3
English 76.1
Latvian 76.5
Kimi K3
OpenRouter, reasoning xhigh main / medium mid / minimal fast EN − LV +7.3
Claude tiered
Opus 4.8 main, Sonnet 4.6 mid, Haiku 4.5 fast EN − LV +3.2
GLM-5.3-Flash
OpenRouter, all tiers EN − LV +4.9
Muse Spark 1.3
OpenRouter, all tiers EN − LV -2.4
Gemini 3.8 Flash
all tiers EN − LV +6.0
DeepSeek V4.1 Flash
OpenRouter, all tiers EN − LV +4.8
OpenAI tiered
GPT-5.6 Sol main, GPT-5.6 Terra mid, GPT-5.4 mini fast EN − LV -0.4

Scores by activity

Each activity is scored 0 to 100 per language.

Stack Requirements extraction Bid analysis Market research
English Latvian English Latvian English Latvian
Kimi K3
92.2 91.5 91.5 93.6 80.0 56.7
Claude tiered
90.2 81.1 95.4 95.0 70.0 70.0
GLM-5.3-Flash
80.3 92.1 95.2 91.2 76.7 54.4
Muse Spark 1.3
90.8 81.6 88.6 95.0 60.0 70.0
Gemini 3.8 Flash
91.4 82.6 100.0 90.9 60.0 60.0
DeepSeek V4.1 Flash
89.7 91.2 87.5 87.6 66.1 50.0
OpenAI tiered
92.2 92.4 82.7 83.6 53.3 53.3

Cost, time and reliability

Cost is the LLM usage of both languages and all activities, including retries. Wall-clock is the duration of the whole cycle. Reliability is completed activity records divided by attempts.

Rank Stack Composite English Latvian EN − LV Total cost Wall-clock Reliability Trials
1
Kimi K3
OpenRouter, reasoning xhigh main / medium mid / minimal fast
84.3 87.9 80.6 +7.3 $46.52 5 h 57 min 100% 1
2
Claude tiered
Opus 4.8 main, Sonnet 4.6 mid, Haiku 4.5 fast
83.6 85.2 82.0 +3.2 $77.69 7 h 7 min 100% 1
3
GLM-5.3-Flash
OpenRouter, all tiers
81.6 84.1 79.2 +4.9 $4.15 13 h 38 min 89% 1
4
Muse Spark 1.3
OpenRouter, all tiers
81.0 79.8 82.2 -2.4 $39.80 2 h 11 min 100% 1
5
Gemini 3.8 Flash
all tiers
80.8 83.8 77.8 +6.0 $70.54 6 h 7 min 100% 1
6
DeepSeek V4.1 Flash
OpenRouter, all tiers
78.7 81.1 76.3 +4.8 $13.90 3 h 14 min 100% 1
7
OpenAI tiered
GPT-5.6 Sol main, GPT-5.6 Terra mid, GPT-5.4 mini fast
76.3 76.1 76.5 -0.4 $30.21 2 h 9 min 100% 1

Stacks, models and judges

Stack Main role Mid role Fast role Judge Run date
Kimi K3
kimi-k3 kimi-k3 kimi-k3 claude-fable-5-1 2026-09-07
Claude tiered
claude-opus-4-8 claude-sonnet-4-6 claude-haiku-4-5 claude-fable-5-1 2026-09-05
GLM-5.3-Flash
glm-5.3-flash glm-5.3-flash glm-5.3-flash claude-fable-5-1 2026-09-07
Muse Spark 1.3
muse-spark-1.3 muse-spark-1.3 muse-spark-1.3 claude-fable-5-1 2026-09-05
Gemini 3.8 Flash
gemini-3.8-flash gemini-3.8-flash gemini-3.8-flash claude-fable-5-1 2026-09-05
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash deepseek/deepseek-v4.1-flash deepseek/deepseek-v4.1-flash gpt-6 2026-09-11
OpenAI tiered
gpt-5.6-sol gpt-5.6-terra gpt-5.4-mini claude-fable-5-1 2026-09-06

Method

A stack is the set of models assigned to the three agent roles: main, mid and fast. Different steps of each agent call different roles. A stack may use one model for all three roles or a different model per role.

The corpus is one municipal school renovation tender with two lots, in parallel English and Latvian versions: regulations, a technical specification, a bill of quantities, a draft contract, a drawing and two bids. One bid is compliant. The other contains 22 planted defects; some appear only in the spreadsheet, some only on scanned certificate pages.

For each stack, language and trial, the benchmark runs requirements extraction on the RFP once, bid analysis once for each of the two bids, and market research on the tender's items once. It uses the same jobs, sandboxes and tools as production, with one exception: every catalogue item is approved automatically so market research can start without a human step.

Bid analysis is scored on planted-defect recall (50%), precision against false alarms on either bid (30%) and verdict correctness (20%); the trial score is the lower of the two bids. Requirements extraction is scored on reference recall (60%), precision (20%) and the lots and criteria structure (20%). Market research is 40% structural (run completed, share of items with at least one quote, share of quotes with a validated source) and 60% a rubric of facts the report must contain. With more than one trial, each activity score is the median across trials. The language composite is the mean of the three activity scores; the overall composite is the mean of the English and Latvian composites.

A separate model reads each output against the answer key in a worksheet; the scorer then applies the fixed weights above. The judge is never one of the models under test, but it may come from the same vendor. The judge for each stack is listed in the stacks table.

Cost is LLM usage only. First-party calls are priced at the provider list price on the run date, with cache reads, cache writes and thinking tokens weighted as the provider bills them. OpenRouter calls use the cost OpenRouter reported. Wall-clock is the time from driver start to finish for both languages, including setup and queue wait.

Reliability is completed activity records divided by the number of attempts; a retry that later succeeds counts as two attempts and one completion. A stack with one trial is a single sample.