A bake-off is a real, dated, reproducible run where the only thing that changes is the model. Same task, same inputs, same prompt. We run one when we need to answer a specific question with money attached — is the cheap model good enough for this job, is our routing actually worth it, can this model be trusted with a number.
Each entry below opens to show what the run found and — the part that matters more — what it does not prove. Every figure is read from the same data file its detail page renders, so nothing here can drift from the run it describes.
Fill a 27-line construction budget from real drawings and a soil study, where inventing a number is worse than leaving the cell blank. Which models can be trusted with a figure that costs money if it is wrong?
A deterministic harness fills the budget from the same drawings for each model, then two independent checks run over the result: a citation check for invented quantities and a scope auditor for figures applied to the wrong extent. The engine validates the arithmetic separately.
Nothing here shows a model can fill correctly when the data is good. The reference run itself could only defensibly fill 1 of 27 lines on most runs, because the drawings lack zone-level areas. Every agreement figure is capped by that ceiling.
For real chatbot turns on live client sites, does picking the model ourselves from measured benchmarks beat handing the choice to OpenRouter’s automatic router — on answer quality and on cost?
Each turn is run through both routers with identical context, then both answers are scored blind by an LLM judge on a 1–5 scale. Cost is the real billed cost of the model each router actually chose.
A sample this size (65 turns) shows the cost gap is real and repeatable across three sites. It is not enough to claim a quality edge — with 56 of 65 diverged turns scoring level, the honest reading is "same answers, far less money", not "better answers".
We use an LLM to judge whether an assistant answered from site data or punted. That judge is an instrument, and running it on a premier model costs real money. Can a cheaper model do the grading without changing the verdict?
Every candidate grades the same probe set with the same rubric and prompt, with only the model varying. Agreement is measured against a fixed premier-model reference, and unparseable responses are counted rather than discarded.
24 probes on one task. This says which judges track our reference on THIS grading job — it does not transfer to a different rubric, and it does not establish that the reference itself is correct. The judge is an instrument; this run calibrates it, it does not validate it.
The useful part of a bake-off is rarely the winner. It is the failure mode you didn't know to look for — a model that fills every cell and gets the arithmetic wrong, a judge that returns text nothing can parse, a router that picks differently every time and lands in the same place. Those are the findings that change what we ship.
So each run above is published with its ceiling attached. A bake-off on one task, with one reference, at one point in time, tells you about that task. We say so rather than letting a narrow result read as a general claim — and when a run is superseded by a fix, the old numbers stay visible, marked, rather than quietly disappearing.