Client bot · value over time
The assistant watches its own real questions, tests whether a less expensive model can answer each kind as well, and proposes the switch — with the evidence attached. Quality is held per question type; the bill comes down. And you approve every change — nothing re-routes on its own.
Measure
Each turn is classed by the question type it matched. The bench re-runs those real questions on candidate models and grades them — so the comparison is your traffic, not a generic benchmark.
Propose
When a less expensive model holds quality on a question type, the loop proposes routing that type to it — with the graded results and the projected saving shown, not a bare recommendation.
Approve
The proposal waits for a human. Approve it and that question type re-tiers live; decline and nothing changes. The premier model stays on the turns that need it.
What you approve
Question type
“hours & location”
128 real turns
Route
Premier → less expensive
same graded quality
Projected saving
≈ 71% on this type
at current volume
Why it's safe
A cheaper route is only ever proposed afterit’s been graded as good as the current one on that exact kind of question — measured on your traffic, not assumed. Question types that genuinely need the premier model keep it. And because a person approves each change, the assistant can never quietly trade your answer quality for a smaller bill.