Routing-efficiency benchmark · fresh split

On problems the models haven't seen, routing beats the frontier for less.

OmnisBench measures how close an LLM routing policy gets to the ideal quality-per-dollar frontier. On old benchmarks the models have seen the answers, so everything scores near 100% and routing looks pointless. On fresh problems, ideal routing hits 93.3%, above the frontier model's 86.7%, at about 60% lower cost. We publish every response, so you can re-grade it offline.

Run it on GitHub See the leaderboard 15 fresh tasks · LiveCodeBench 2025+ 4-model pool run 2026-08-20 Apache-2.0
93.3%
fresh-task success under ideal routing (oracle)
~60%
cheaper than the frontier model, at higher quality
60%
the cheap model alone on fresh tasks, down from 90% contaminated
100%
of the leaderboard re-derivable from results.json
The thesis

Every policy on the fresh split

Up and to the left is better: more task success, less money. This is the fresh split, the LiveCodeBench problems the models can't have memorised. The two policies on the dashed Pareto frontier aren't beaten on both axes by anything else. Ideal routing (oracle) sits top-left; always calling the frontier model buys less quality for more than twice the spend.

Cost per 1,000 requests (log-free, linear USD) vs. task success. Points on the accent-coloured frontier are non-dominated; muted points are beaten on both axes. Hover any point for detail.
The scoreboard

Leaderboard

Four policies over the same 15 fresh LiveCodeBench tasks and the same pinned prices. oracle is the theoretical ceiling: the cheapest model that did solve each item, chosen with hindsight. It's the target a real router aims at, not one you could ship.

PolicyWhat it does Task success$ / 1kvs frontier modelFrontier

The honest read: on old benchmarks the cheap model looks nearly perfect, but that's contamination. On these fresh tasks gpt-5-nano alone scores 60%, ideal routing recovers that to 93.3%, and it does so for about 60% less than always calling the frontier model. That gap is the prize, and it only shows up once the data is clean.

Under the hood

Where ideal routing sends the work

For each task, oracle picks the cheapest model that solves it. On the fresh set it solves 9 of 15 most cheaply with the tiniest model and escalates to a stronger one for the hard six. That escalation is the 20% frontier-escape rate, and it's where the quality comes from.

The contamination effect

Same policies, scored on the likely-contaminated split (HumanEval + GSM8K) and the fresh split (LiveCodeBench 2025+).

PolicyLikely-contaminated (20)Fresh, 2025+ (15)
cheapest model only (gpt-5-nano)90.0%60.0%
oracle (ideal routing)100.0%93.3%
always the frontier (claude-opus-5)100.0%86.7%

On the contaminated split the models have seen the problems, so everything lands near 100% and routing looks like it saves nothing. That flatness is the artefact. The fresh column is the real picture.

Measuring contamination, not guessing it

The perturbation-gap probe

A "fresh" tag is a judgement. This probe makes it a number. Score each model twice on the same problems, once on the original and once on a reworded copy that keeps the test cases identical, then publish the drop. A model that relied on the exact wording scores lower when you change it, so a large gap flags contamination on that set for that model. A gap near zero means the score was earned. You can re-grade it offline, the same as everything else here.

ModelGap, reworded by claude-sonnet-5 (n=15)Gap, reworded by gpt-5 (n=17)
claude-opus-5+20.0pp+5.9pp
claude-haiku-4-5+40.0pp+5.9pp
gpt-5+20.0pp+23.5pp
gpt-5-nano+20.0pp+52.9pp

Two rewriters, one from each provider, reword the whole pool, so no gap rides on same-family phrasing and you can see how much the rewriter itself moves the number. One result survives that test: gpt-5 gives up about 20 points whoever rewords it, a real sign its score here leaned on the exact phrasing. The rest swing with the rewriter. The Anthropic models barely move under gpt-5 and fall hard under sonnet, gpt-5-nano the other way, so at 15 to 17 problems the rewriter and the surviving tasks move the number more than the model does. What holds for all of them: every model loses at least 6 points when a fresh problem is reworded. A post-cutoff date keeps the exact problem out of training, the wording and structure are not, so read the fresh tag as necessary and not the whole story. Method, both runs and the reworded prompts: docs/perturbation-probe.md.

Why you can trust it

Don't take our word for it, re-grade it

Most routing savings claims are unverifiable marketing. OmnisBench publishes the exact model response for every task inside results.json. omnisbench verify re-runs the graders against those responses and re-derives this entire leaderboard, quality and cost, with no API calls and no keys. Change one stored answer and it fails.

# reproduce the whole run, or audit the published one
$ pip install -e .
$ pip install datasets                       # for the LiveCodeBench fresh split
$ python -m omnisbench.cli run    --config configs/fresh-run.yaml --run runs/mine
$ python -m omnisbench.cli verify --run runs/fresh-16k-2026-08-20  # zero-API re-grade
VERIFY OK
Questions

Frequently asked

Straight answers to the questions people search for about LLM routing and OmnisBench.

What is an LLM router?
An LLM router inspects each request and sends it to the cheapest model that can still handle it, instead of routing everything to one expensive frontier model. OmnisBench measures how close a given routing policy gets to that ideal, which is the most task success per dollar.
How much can LLM routing save?
It depends how contaminated your benchmark is. On old suites the models have seen the answers and every policy scores near 100%, so routing looks pointless. On a fresh split of LiveCodeBench problems published after the training cutoff, the cheapest model drops to 60%, ideal routing reaches 93.3% (above the frontier model's 86.7%) at about 60% lower cost, and routing recovers thirty-three points the contaminated benchmarks were hiding. Every figure re-grades from the published results.
Is OmnisBench an open-source alternative to Weave Router or RouterBench?
OmnisBench is an open, Apache-2.0 benchmark for LLM routing efficiency. It's not a router itself. Where vendors publish savings claims you can't check, OmnisBench publishes the model response for every task and lets anyone re-grade the numbers offline with a single command. It complements the academic RouterBench by staying live and cost-current, and by letting anyone re-verify it.
How do I verify the results myself?
Download results.json and run omnisbench verify. It re-runs the graders against the published responses and re-derives the full leaderboard (quality and cost) with no API calls and no keys. Change one stored answer and it fails.
Which models and datasets does it cover?
The headline run scores four models (Claude Opus 5, GPT-5, Claude Haiku 4.5, GPT-5-nano) on two splits: a fresh set of LiveCodeBench problems from 2025 onward, and a likely-contaminated set of HumanEval and GSM8K. More datasets and models are config-only additions.
Read the numbers honestly

Caveats, stated up front

oracle is a ceiling, not a product

It's chosen with hindsight (cheapest model that did solve each task). No live router can be this good; the gap between a real router and this line is the real scorecard.

Small samples

15 fresh tasks and 20 contaminated. It's a pilot that shows the method and the direction, not a final verdict. Widening the fresh set is the roadmap.

Difficulty is tangled with freshness

The fresh set is competitive-programming, which is both newer and harder than grade-school maths. Some of the drop is difficulty, not contamination on its own.

Pinned, dated prices

Costs come from a committed pricing snapshot (2026-08-18), not live lookups, so the numbers stay reproducible even as list prices move.