a grounded support turn: given a website's own pages, answer the visitor's question from those pages or refuse — never fill a gap from general knowledge
8 models · 15 turns · run August 9, 2026
openai/gpt-oss-120b scored 12/12 on refusals and 14/14 on grounded answers across 2 independent runs — identical to the claude-sonnet-5 baseline — at $1.15 per 1,000 turns against $31.23, a 27× saving, and 4.5× faster.
The negative controls refused to fail. 2 models were included specifically because we expected them to underperform — the smallest and cheapest options available. They scored perfectly on first contact.
Most benchmarks score whether a model produced an answer. Here six of the fifteen turns are questions the source CANNOT answer, and the only correct response is to decline. A confident, fluent, factually-true answer on one of those rows is a FAILURE: the model cannot have grounded it, so it came from priors, and the next one will be wrong with equal confidence.
Six of eight models scored perfectly, including both models chosen as negative controls because we expected them to fail. That is the finding, and it cuts both ways: on this corpus at this prompt, the task no longer separates an 8B model from a frontier one, so these numbers should not be read as a general capability ranking. They answer one narrow question — can a cheaper model do THIS job — and the honest answer is that a harder fixture is now needed to tell the top of the field apart.
A candidate replaces the baseline only by matching it on BOTH axes. Fabricating once disqualifies, regardless of price or speed — for the kind of site this was run against, an invented figure costs real money downstream.
| Model | Runs | Refused correctly | Grounded correctly | Errors | Median | $ / 1k turns |
|---|---|---|---|---|---|---|
| claude-sonnet-5Baselineanthropic | 2 | 12/12 | 14/14 | 0 | 6.7s | $31.23 |
| llama-3.1-8b-instantControlgroq | 1 | 6/6 | 7/7 | 0 | 0.8s | $0.29 |
| openai/gpt-oss-20bControlgroq | 1 | 6/6 | 7/7 | 0 | 0.9s | $0.85 |
| openai/gpt-oss-120bgroq | 2 | 12/12 | 14/14 | 0 | 1.5s | $1.15 |
| deepseek-v3.2openrouter | 2 | 12/12 | 14/14 | 0 | 4.1s | $1.61 |
| gemini-3-flash-previewgoogle | 2 | 12/12 | 14/14 | 0 | 4.8s | $3.16 |
| gemma-4-31b-itgoogle | 2 | 3/4 | — | 26 | 10.8s | — |
| qwen3-32bopenrouter | 2 | 8/12 | 14/14 | 0 | 13.9s | $0.60 |
gemma-4-31b-it is reported as untested, not failed — its free tier exhausted its quota mid-run. the call failed (quota, unreachable provider). EXCLUDED from every rate — an unreachable model is unknown, never a pass and never a failure.
Answered 4 of 12 questions the source could not support. It also declines and then answers in the same reply — which is why a scorer that matches refusal phrases rather than the decision itself would mark it as correct.
6 of the 15 turns cannot be answered from the source, so declining is the only correct response. 7 are answerable, and the rest test whether the assistant remembers the previous turn.
Run against a live tenant's corpus. No client name, tenant identifier, page address, source text or reply body is published here — only the shape of each turn and whether the model answered or refused correctly.
Each of these produced a confident, wrong conclusion that had to be caught by reading the replies by hand. Two of them put the wrong model at the top of an earlier draft of this page.