The open, reproducible LLM routing benchmark.
OmnisBench measures how close a routing policy gets to the ideal quality-per-dollar frontier. We publish every response, so anyone can re-grade the numbers offline with no API calls. It's a benchmark, not a router: it grades routing policies, it doesn't route your traffic.
TL;DR: Routers make savings claims that are hard to check. OmnisBench is the open scoreboard that checks them: it plots every routing policy against an ideal quality-per-dollar frontier over a fixed task suite, publishes the underlying model responses, and lets anyone re-grade the entire leaderboard offline with omnisbench verify. On a fresh split the models can't have memorised, ideal routing hit 93.3% task success, above the frontier model itself, at about 60% lower cost.
Routing efficiency, not routing product
OmnisBench doesn't ship a router and doesn't touch your traffic. It runs a fixed suite of tasks against a fixed model pool under several routing policies (including oracle, the theoretical ceiling that always picks the cheapest model that actually solved a given task) and plots each policy's task success against its cost. The policies on the resulting Pareto frontier aren't beaten on both axes by anything else in the pool. That frontier is the yardstick any real router, including a future one from us, gets graded against.
Routing beats the frontier model on fresh tasks
On a fresh split of 15 LiveCodeBench problems published in 2025 or later, so the models can't have trained on them, ideal routing reached 93.3% task success at $53 per 1,000 requests. That's above the frontier model itself, Claude Opus 5, which managed 86.7% at $138/1k, and it costs less than half. No single model solves every fresh problem, so routing to the best one per task wins on quality and price at once. The cheapest model alone drops to 60% on fresh work, down from 90% on the older suites, which is how much those contaminated benchmarks were flattering it. The honest caveat: 15 tasks is a pilot that shows the method and the direction, not a final verdict.
| Policy | What it does | Task success | $ / 1k | Frontier |
|---|---|---|---|---|
| oracle | cheapest model that solved each task | 93.3% | $53.30 | frontier |
| always_cheap | always gpt-5-nano | 60.0% | $4.10 | frontier |
| always_big | always claude-opus-5 (frontier model) | 86.7% | $138.10 | dominated |
Anyone can check the number, not just trust it
Most LLM routing savings claims are marketing you can't verify: a percentage on a landing page with no published methodology, no model, no dataset. OmnisBench takes the opposite approach: the code is Apache-2.0, the task suite is named and fixed, the model pool is disclosed, and every individual model response is published inside results.json. Running omnisbench verify re-runs the graders against those stored responses and re-derives the entire leaderboard, quality and cost, with zero API calls and no keys. Change one stored answer and it fails.
Run it yourself
The full run and the offline verification live in the public GitHub repository. No account, no API keys required to re-grade the published results.
# reproduce the whole run, or just audit the published one $ pip install -e . $ python scripts/prepare_datasets.py # HumanEval + GSM8K → data/ $ python -m omnisbench.cli run --config configs/fresh-run.yaml --run runs/mine $ python -m omnisbench.cli verify --run runs/fresh-16k-2026-08-20 # zero-API re-grade VERIFY OK
See how it stacks up
The fresh split →
Why contamination flatters the old benchmarks, and how routing earns more on tasks the models can't have seen.
OmnisBench vs RouterBench →
How a live, one-command-reproducible benchmark compares to the respected academic RouterBench baseline.
Weave Router alternative →
Weave Router is a closed-routing product with unverified savings claims. Here's how to check them.
Want the full results, charts, and leaderboard? See the homepage →