On problems the models haven't seen, routing beats the frontier for less.
OmnisBench measures how close an LLM routing policy gets to the ideal quality-per-dollar frontier. On old benchmarks the models have seen the answers, so everything scores near 100% and routing looks pointless. On fresh problems, ideal routing hits 93.3%, above the frontier model's 86.7%, at about 60% lower cost. We publish every response, so you can re-grade it offline.
Every policy on the fresh split
Up and to the left is better: more task success, less money. This is the fresh split, the LiveCodeBench problems the models can't have memorised. The two policies on the dashed Pareto frontier aren't beaten on both axes by anything else. Ideal routing (oracle) sits top-left; always calling the frontier model buys less quality for more than twice the spend.
Leaderboard
Four policies over the same 15 fresh LiveCodeBench tasks and the same pinned prices. oracle is the theoretical ceiling: the cheapest model that did solve each item, chosen with hindsight. It's the target a real router aims at, not one you could ship.
| Policy | What it does | Task success | $ / 1k | vs frontier model | Frontier |
|---|
The honest read: on old benchmarks the cheap model looks nearly perfect, but that's contamination. On these fresh tasks gpt-5-nano alone scores 60%, ideal routing recovers that to 93.3%, and it does so for about 60% less than always calling the frontier model. That gap is the prize, and it only shows up once the data is clean.
Where ideal routing sends the work
For each task, oracle picks the cheapest model that solves it. On the fresh set it solves 9 of 15 most cheaply with the tiniest model and escalates to a stronger one for the hard six. That escalation is the 20% frontier-escape rate, and it's where the quality comes from.
The contamination effect
Same policies, scored on the likely-contaminated split (HumanEval + GSM8K) and the fresh split (LiveCodeBench 2025+).
| Policy | Likely-contaminated (20) | Fresh, 2025+ (15) |
|---|---|---|
| cheapest model only (gpt-5-nano) | 90.0% | 60.0% |
| oracle (ideal routing) | 100.0% | 93.3% |
| always the frontier (claude-opus-5) | 100.0% | 86.7% |
On the contaminated split the models have seen the problems, so everything lands near 100% and routing looks like it saves nothing. That flatness is the artefact. The fresh column is the real picture.
The perturbation-gap probe
A "fresh" tag is a judgement. This probe makes it a number. Score each model twice on the same problems, once on the original and once on a reworded copy that keeps the test cases identical, then publish the drop. A model that relied on the exact wording scores lower when you change it, so a large gap flags contamination on that set for that model. A gap near zero means the score was earned. You can re-grade it offline, the same as everything else here.
| Model | Gap, reworded by claude-sonnet-5 (n=15) | Gap, reworded by gpt-5 (n=17) |
|---|---|---|
| claude-opus-5 | +20.0pp | +5.9pp |
| claude-haiku-4-5 | +40.0pp | +5.9pp |
| gpt-5 | +20.0pp | +23.5pp |
| gpt-5-nano | +20.0pp | +52.9pp |
Two rewriters, one from each provider, reword the whole pool, so no gap rides on same-family phrasing and you can see how much the rewriter itself moves the number. One result survives that test: gpt-5 gives up about 20 points whoever rewords it, a real sign its score here leaned on the exact phrasing. The rest swing with the rewriter. The Anthropic models barely move under gpt-5 and fall hard under sonnet, gpt-5-nano the other way, so at 15 to 17 problems the rewriter and the surviving tasks move the number more than the model does. What holds for all of them: every model loses at least 6 points when a fresh problem is reworded. A post-cutoff date keeps the exact problem out of training, the wording and structure are not, so read the fresh tag as necessary and not the whole story. Method, both runs and the reworded prompts: docs/perturbation-probe.md.
Don't take our word for it, re-grade it
Most routing savings claims are unverifiable marketing. OmnisBench publishes the exact model response for every task inside results.json. omnisbench verify re-runs the graders against those responses and re-derives this entire leaderboard, quality and cost, with no API calls and no keys. Change one stored answer and it fails.
# reproduce the whole run, or audit the published one $ pip install -e . $ pip install datasets # for the LiveCodeBench fresh split $ python -m omnisbench.cli run --config configs/fresh-run.yaml --run runs/mine $ python -m omnisbench.cli verify --run runs/fresh-16k-2026-08-20 # zero-API re-grade VERIFY OK
Frequently asked
Straight answers to the questions people search for about LLM routing and OmnisBench.
What is an LLM router?
How much can LLM routing save?
Is OmnisBench an open-source alternative to Weave Router or RouterBench?
How do I verify the results myself?
Which models and datasets does it cover?
Caveats, stated up front
oracle is a ceiling, not a product
It's chosen with hindsight (cheapest model that did solve each task). No live router can be this good; the gap between a real router and this line is the real scorecard.
Small samples
15 fresh tasks and 20 contaminated. It's a pilot that shows the method and the direction, not a final verdict. Widening the fresh set is the roadmap.
Difficulty is tangled with freshness
The fresh set is competitive-programming, which is both newer and harder than grade-school maths. Some of the drop is difficulty, not contamination on its own.
Pinned, dated prices
Costs come from a committed pricing snapshot (2026-08-18), not live lookups, so the numbers stay reproducible even as list prices move.