Routing-efficiency benchmark · fresh split

The open, reproducible LLM routing benchmark.

OmnisBench measures how close a routing policy gets to the ideal quality-per-dollar frontier. We publish every response, so anyone can re-grade the numbers offline with no API calls. It's a benchmark, not a router: it grades routing policies, it doesn't route your traffic.

15 fresh tasks · LiveCodeBench 2025+ 4-model pool Apache-2.0 omnisbench verify · zero API calls

TL;DR: Routers make savings claims that are hard to check. OmnisBench is the open scoreboard that checks them: it plots every routing policy against an ideal quality-per-dollar frontier over a fixed task suite, publishes the underlying model responses, and lets anyone re-grade the entire leaderboard offline with omnisbench verify. On a fresh split the models can't have memorised, ideal routing hit 93.3% task success, above the frontier model itself, at about 60% lower cost.

What it measures

Routing efficiency, not routing product

OmnisBench doesn't ship a router and doesn't touch your traffic. It runs a fixed suite of tasks against a fixed model pool under several routing policies (including oracle, the theoretical ceiling that always picks the cheapest model that actually solved a given task) and plots each policy's task success against its cost. The policies on the resulting Pareto frontier aren't beaten on both axes by anything else in the pool. That frontier is the yardstick any real router, including a future one from us, gets graded against.

The fresh result

Routing beats the frontier model on fresh tasks

On a fresh split of 15 LiveCodeBench problems published in 2025 or later, so the models can't have trained on them, ideal routing reached 93.3% task success at $53 per 1,000 requests. That's above the frontier model itself, Claude Opus 5, which managed 86.7% at $138/1k, and it costs less than half. No single model solves every fresh problem, so routing to the best one per task wins on quality and price at once. The cheapest model alone drops to 60% on fresh work, down from 90% on the older suites, which is how much those contaminated benchmarks were flattering it. The honest caveat: 15 tasks is a pilot that shows the method and the direction, not a final verdict.

PolicyWhat it does Task success$ / 1kFrontier
oracle cheapest model that solved each task 93.3% $53.30 frontier
always_cheap always gpt-5-nano 60.0% $4.10 frontier
always_big always claude-opus-5 (frontier model) 86.7% $138.10 dominated
Why open + reproducible matters

Anyone can check the number, not just trust it

Most LLM routing savings claims are marketing you can't verify: a percentage on a landing page with no published methodology, no model, no dataset. OmnisBench takes the opposite approach: the code is Apache-2.0, the task suite is named and fixed, the model pool is disclosed, and every individual model response is published inside results.json. Running omnisbench verify re-runs the graders against those stored responses and re-derives the entire leaderboard, quality and cost, with zero API calls and no keys. Change one stored answer and it fails.

Reproduce it

Run it yourself

The full run and the offline verification live in the public GitHub repository. No account, no API keys required to re-grade the published results.

# reproduce the whole run, or just audit the published one
$ pip install -e .
$ python scripts/prepare_datasets.py       # HumanEval + GSM8K → data/
$ python -m omnisbench.cli run    --config configs/fresh-run.yaml --run runs/mine
$ python -m omnisbench.cli verify --run runs/fresh-16k-2026-08-20  # zero-API re-grade
VERIFY OK

github.com/Fortitude-Group/OmnisBench →

Related reading

See how it stacks up

The fresh split →

Why contamination flatters the old benchmarks, and how routing earns more on tasks the models can't have seen.

OmnisBench vs RouterBench →

How a live, one-command-reproducible benchmark compares to the respected academic RouterBench baseline.

Weave Router alternative →

Weave Router is a closed-routing product with unverified savings claims. Here's how to check them.

Want the full results, charts, and leaderboard? See the homepage →

Questions

Frequently asked

What is an LLM routing benchmark?
An LLM routing benchmark measures how well a routing policy sends each request to the cheapest model that can still handle it, instead of scoring the router product itself. It compares policies against an ideal quality-per-dollar frontier so you can see how much of the theoretical savings a real router actually captures.
Is OmnisBench really reproducible?
Yes. OmnisBench publishes the exact model response for every task in results.json, and the omnisbench verify command re-runs the graders against those published responses and re-derives the entire leaderboard offline, with zero API calls and no keys. Change one stored answer and verification fails.
What does OmnisBench cover?
The headline run scores four models (Claude Opus 5, GPT-5, Claude Haiku 4.5, GPT-5-nano) on two splits: a fresh set of 15 LiveCodeBench problems from 2025 onward, which the models can't have trained on, and a likely-contaminated set of HumanEval and GSM8K. It's Apache-2.0 licensed, and more datasets and models are config-only additions.