Does routing hold on tasks the models haven't seen?
HumanEval and GSM8K are old, and the models have almost certainly read the answers. So we built a fresh split: the same routing policies, scored on problems published after the models could have trained on them. The short version is that contamination was real, routing earns more on fresh work, and we caught our own harness quietly strangling the models on the way. Every number here re-grades offline.
Contaminated versus fresh
Same four policies, scored separately on the old benchmarks (HumanEval + GSM8K) and on a fresh set of LiveCodeBench problems dated 2025 or later. oracle is ideal routing, the cheapest model that did solve each task.
| Policy | Likely-contaminated (20) | Fresh, 2025+ (15) |
|---|---|---|
| cheapest model only (gpt-5-nano) | 90.0% | 60.0% |
| oracle (ideal routing) | 100.0% | 93.3% |
| always the frontier model (claude-opus-5) | 100.0% | 86.7% |
| how often ideal routing reached for the frontier | 0% | 20% |
Three things fall out of that. The frontier model's fine on fresh problems (86.7%). The cheapest model drops from 90% to 60%, so the old benchmarks were flattering it, which is the contamination the commenters warned about. And routing earns far more on the fresh set: on the old tasks it buys ten points over the cheap model, on the fresh tasks it buys thirty-three, while escalating to the expensive model a fifth of the time instead of never. Routing pays off hardest exactly where the models are tested on work they can't have memorised.
We caught the harness truncating the models
The first fresh run looked catastrophic, and it was wrong. Our output budget was 4,096 tokens, which is plenty for "write this short function" and nowhere near enough for a hard competitive-programming problem. The reasoning models thought at length, hit the ceiling mid-thought, and emitted no code at all. Every one of the frontier model's fresh failures was a blank page. We only spotted it because we publish every response and could read them. We made the budget configurable, gave the hard suite room, and re-ran. Truncation dropped from every failure to a single task. A closed benchmark makes that same mistake and ships the number, and nobody's the wiser.
The dates behind the tag
Calling a split "fresh" is a judgement, so here's the evidence. The 15 fresh tasks are LiveCodeBench problems published between 2025-01-04 and 2025-04-06, median 2025-02-22, filtered to anything dated 2025-01-01 or later. The model providers don't publish exact training cutoffs, so we hand you the release dates and let you judge how far past the line this run sits, rather than a bare clean-or-contaminated call. Those dates ship in results.json and re-derive under omnisbench verify like everything else. The measured version is the perturbation-gap probe, described step by step below.
How the perturbation-gap test works
- Take a fresh problem and its exact test cases, the input-to-output pairs from LiveCodeBench.
- A separate rewriting model, never the one being graded on that problem, rewrites the statement. It changes the prose, the variable names and the surface story, and holds the input format, the output format and every constraint word for word. The test cases never change, so the correct answer is identical.
- Score every model twice, on the original and on the reworded copy. Grading runs the model's code against the same unchanged tests, so a pass means the same thing both times.
- Drop any reworded problem no model can solve, a sign the rewrite broke it, and any problem no model solved at all.
- The gap is the drop, original accuracy minus reworded accuracy. A model that leaned on the exact wording falls. A model that understood the problem holds.
The pool spans two providers, and a rewriter can favour or punish its own family's style, so we reword the whole pool twice, once with claude-sonnet-5 and once with gpt-5, and read each model's gap under each. That shows how much the rewriter itself moves the number, not the model.
| Model | Gap, reworded by claude-sonnet-5 (n=15) | Gap, reworded by gpt-5 (n=17) |
|---|---|---|
| claude-opus-5 | +20.0pp | +5.9pp |
| claude-haiku-4-5 | +40.0pp | +5.9pp |
| gpt-5 | +20.0pp | +23.5pp |
| gpt-5-nano | +20.0pp | +52.9pp |
One number survives changing the rewriter: gpt-5 gives up about 20 points whoever rewords it, a real sign its score on these recent problems leaned on the phrasing. Every other gap swings with the rewriter. The Anthropic models barely move under gpt-5 and fall hard under sonnet, gpt-5-nano the other way, so at 15 to 17 problems the rewriter and the surviving tasks move the number more than the model does. What holds for all four: every model loses at least 6 points when a fresh problem is reworded, so a post-cutoff date is necessary and not the whole story. Every reworded prompt and response is published, and omnisbench verify re-derives the gaps offline. Both runs and the method: docs/perturbation-probe.md.
Caveats, up front
Small samples
20 contaminated and 15 fresh tasks. It's a first signal, not a verdict, and the suite grows from here.
Difficulty is tangled with freshness
The fresh set is competitive-programming, which is both newer and harder than grade-school maths. Part of the drop is difficulty, not contamination on its own. We can't fully separate the two yet.
One residual truncation
Even at 16k tokens, one fresh problem still ran the model out of room. That one's genuine difficulty now, not a harness fault.
oracle is a ceiling
It's chosen with hindsight. A shippable router can't be this good. The gap to it is the real scorecard.
Re-grade the run offline
The whole run, responses included, is in the repo. verify re-runs the graders against the published answers and re-derives this table with no API calls.
# re-grade the corrected fresh run, no keys needed $ pip install -e . $ python -m omnisbench.cli verify --run runs/fresh-16k-2026-08-20 VERIFY OK
See the run and the code on GitHub →
Back to the full leaderboard →