Switchyard — the smart router that learns. Groundyard — your new OS into your site and data. We stress test models every night and publish the results. Decide on value.
Building an AI product? We measure your real tasks and hand back a routing table — the least-expensive model that still passes, savings included. Read our principles →
What we're building
Most AI products ask you to take their word for it. Here is what we think is wrong with that, what we're building instead, and — plainly — which parts are finished. Read the long version →
Every AI vendor's numbers come from the vendor. We run our own stress tests nightly across every major provider — deterministic validators, no LLM grading another LLM — and publish the whole thing free. You decide what's true.
We'd rather the assistant decline than invent. On our own traffic Groundyard declined 42 of 145 grounded questions rather than guess. Small numbers so far — and we'd rather publish them than hide them.
Nothing re-tiers your assistant without a person saying yes. Every answer can be marked wrong by the person reading it, with a reason. Spend caps are checked before the model is ever called.
Switchyard already measures your assistant’s real questions against cheaper models and proposes a better-value route with the evidence attached. Today an operator approves each one. Next: you get that dial yourself, inside bounds you set.
Every performance and cost claim on this site, with the method, the date, the command to reproduce it — and what each number does not cover.
The tools and models we investigated and declined, with the reason each time. Stars measure hype, not fit — so we publish the misses too.
The products
Four products, one measured foundation. The bench below is not for sale — it is the proof the other three stand on.
The measured routing gateway — every request to the least-expensive model that still passes. Measured 7–10× cheaper than routing on autopilot.
The answer engine built from your website — grounded in your data, bilingual, voice-capable, behind hard spend caps. It answers from your sources or it abstains.
The proof
Every model on every test, with cost per task, latency, and pass rate. Free, no signup.
The 3 models not dominated by any lower-cost model on pass rate. Pick from this list and you're on the efficient frontier.
Full analysisMethodology
8 deterministic tests per model. Pass means the response was correct — we execute code, parse JSON, check facts. Not just “model returned 200.”
Basic math, multi-step logic
Python function synthesis with sandbox execution
Strict JSON schema with type validation
BBH-lite: boolean + web-of-lies + counting
Find a code buried in 3.5k tokens of filler
Fact extraction from in-prompt data
Function/tool invocation with correct arguments
Explore
What's happening
7 provider incidents in the last 48h — details on the News page.
Beyond the leaderboard
We don't stop at which model. The same measured, cost-first lens runs across the whole stack — prompt routing, caching and caps, local tools that shrink the input before a metered call, and self-hosting where the math works. Real levers, run on our own product first, with the numbers to prove each one.
Four levers that actually cut the bill — routing, caching, self-hosting, the closed loop. A measured 31% saving per turn.
MarkItDown + Graphify: local $0 tools that shrink documents and codebases before the model ever reads them.
Prompt compression, agent memory, and cost observability — open-source, rated for a cost-first stack.
Every model on every test, with cost per task, latency, throughput, and pass rate — free, no signup, refreshed nightly. Or let us benchmark your real tasks and hand back the routing table.
Your assistant and its live business data inside Claude Desktop — the AI you already pay for does the talking; we serve the data.
per 1,000 successful task runs
FREE model with 100% pass rate — strong default for cost-sensitive workloads
How to Evaluate LLM Provider Performance Across Latency, Throughput, and Uptime
OpenRouter published a technical guide on measuring and comparing LLM provider performance across multiple dimensions. The piece explains that the same model produces different latency, throughput, and uptime characteristics depending on the provider's infrastructure, quantization choices, and load-handling strategies. Developers can use concrete measurement techniques to evaluate providers and build routing policies that direct requests to the best-performing endpoint for their use case, accounting for the fact that provider differences materially affect real-world behavior beyond just model selection.
IFEval-style: bullet format + keyword constraints
The latest important releases, breakthroughs, and new hardware — aggregated from provider blogs and curated AI-news outlets, classified by Claude. 7 recent incidents
The Token Tax: why a metered AI bill grows even when your usage doesn’t — and how to defuse it.