When a wrong number costs money, silence is the right answer.
Construction quantity takeoff — fill a 27-line Mexican budget from DXF drawings and a soil study, leaving a cell blank rather than inventing a figure. Most benchmarks reward a model for producing an answer. This one doesn’t: a blank cell is frequently correct, because the drawings genuinely don’t support a figure. What we measure is whether the numbers a model does give can be trusted.
Results generated 2026-07-26 from the run log. This page reads the same JSON the harness writes, so it cannot drift from what was actually measured.
The three ways to be wrong
Invented
a quantity with NO cited source — a fabricated number
Scope violation
a REAL, correctly-measured figure applied to the WRONG EXTENT (e.g. the site boundary used for a zone-limited concept). More dangerous than an invented value: the citation is genuine, so every other check passes.
Validator rejected
the engine's own validator refused the answer (e.g. amount != quantity x unit_price). DISQUALIFYING and invisible in every other field: such a run can show the highest `filled` with invented=0 and scope_violations=0.
The verdict rule
any invented value, scope violation, or validator rejection disqualifies, regardless of cost or fill count
Read this before the table
Never rank by fill rate.Blanks are frequently the correct answer here. The proof is in the table below: the candidate that filled the most lines was rejected by the engine’s own validator for bad arithmetic, while showing zero invented values and zero scope violations. Sorting by “filled” would put the worst answer on top.
A clean sheet is not a win. Under the rule above, a model that fills nothing cannot fail — and one candidate filled 0 of 27 lines. That is abstention: untested, not safe. Clean means the figures it gave can be trusted, not that it did the job.
Agreement is only as strong as the reference. The reference run could defensibly fill just 1 of 27 lines, because the drawings lack zone-level areas. A high agreement score here largely measures agreement about staying blank — it does not show a model can fill correctly when the data is good.
This is a small sample. 3 bake-offs and 6 reference runs on one budget. It is enough to show a failure mode exists; it is not enough to rank models.
How to read these numbers — from the run itself
Do not rank by `filled`. Disqualify first on invented / scope_violations / validator_rejected, and only then compare cost and latency among survivors. A candidate with filled=0 has not passed — it abstained, which is untested rather than safe. Agreement figures are only as strong as the reference run: check the matching entry in `reference_runs` for how many lines the reference itself filled before trusting a high agreement score.
What we concluded
We stayed on our existing stack. Of the candidates, one was rejected by the validator, one committed a real scope violation, and one abstained entirely. Gemini 3.6 Flash matched the reference exactly at roughly a tenth of the cost and is the one we are watching — but nothing here certifies it. It agreed with a reference that could only fill 1 of 27 lines, and it produced fewer findings than that reference did. No clean multi-filled reference exists yet, so this run cannot show that any model fills correctly when the drawings do support a figure.
27
budget lines
per run
2/4
clean sheets
nothing invented, misapplied or refused
3/4
filled anything
gave at least one figure
The runs
bakeoff 2026-07-26 17:55
current
experiment e1d83298
Model
Filled
Invented
Scope
Rejected
Findings
Out
Latency
gemini-3.6-flash
gemini
1/27
0
0
—
6
4.2k
46.7s
gpt-5.4-2026-03-05
openai
5/27
0
0
yes
27
5.6k
41.0s
answer refused
llama-3.3-70b-versatile
groq
0/27
0
0
—
3
2.9k
5.7s
filled nothing
kimi-for-coding
kimi-cli · subscription
4/27
0
3
—
13
null
110.7s
bakeoff 2026-07-26 12:32
superseded
experiment 0434d68f
Kept for the record. This run predates a fix to the scope auditor, which was flagging one concept that the rules genuinely permit to use the site boundary — so scope violations shown here may not be real. Read the current run above.
Model
Filled
Invented
Scope
Rejected
Findings
Out
Latency
gemini-3.6-flash
gemini
1/27
0
0
—
23
6.0k
51.5s
gpt-5.4-2026-03-05
openai
4/27
0
3
—
19
4.5k
31.0s
llama-3.3-70b-versatile
groq
0/27
0
0
—
29
4.5k
8.4s
filled nothing
bakeoff 2026-07-26 12:11
superseded
experiment 15dce07e
Kept for the record. This run predates a fix to the scope auditor, which was flagging one concept that the rules genuinely permit to use the site boundary — so scope violations shown here may not be real. Read the current run above.
Model
Filled
Invented
Scope
Rejected
Findings
Out
Latency
kimi-for-coding
kimi-cli · subscription
4/27
0
3
—
12
null
101.5s
Reference runs
The baseline the candidates are measured against — the same harness run on our own stack. Included because a bake-off without a control is just a list of numbers.
These reference runs predate the scope-auditor fix and have not been re-run. One shows three scope violations that may be the same false positive corrected in the current bake-off — we have not verified it either way, so it is left as measured. Note also the spread: most reference runs could defensibly fill only 1 of 27 lines, which is the ceiling every agreement figure on this page is measured against.
Model
Filled
Invented
Scope
Rejected
Findings
Out
Latency
claude-opus-5
agent-sdk · subscription
1/27
0
0
—
32
10.4k
—
claude-opus-5
agent-sdk · subscription
1/27
0
0
—
12
8.1k
—
claude-opus-5
agent-sdk · subscription
1/27
0
0
—
35
13.4k
—
claude-opus-5
agent-sdk · subscription
1/27
0
0
—
15
10.6k
—
claude-opus-5
agent-sdk · subscription
5/27
0
3
—
14
11.4k
—
claude-opus-5
agent-sdk · subscription
1/27
0
0
—
15
8.3k
—
Use the data
Every figure on this page comes from one JSON file, served openly so you can check it or chart it yourself.