The fresh split

Does routing hold on tasks the models haven't seen?

HumanEval and GSM8K are old, and the models have almost certainly read the answers. So we built a fresh split: the same routing policies, scored on problems published after the models could have trained on them. The short version is that contamination was real, routing earns more on fresh work, and we caught our own harness quietly strangling the models on the way. Every number here re-grades offline.

The result

Contaminated versus fresh

Same four policies, scored separately on the old benchmarks (HumanEval + GSM8K) and on a fresh set of LiveCodeBench problems dated 2025 or later. oracle is ideal routing, the cheapest model that did solve each task.

PolicyLikely-contaminated (20)Fresh, 2025+ (15)
cheapest model only (gpt-5-nano)90.0%60.0%
oracle (ideal routing)100.0%93.3%
always the frontier model (claude-opus-5)100.0%86.7%
how often ideal routing reached for the frontier0%20%

Three things fall out of that. The frontier model's fine on fresh problems (86.7%). The cheapest model drops from 90% to 60%, so the old benchmarks were flattering it, which is the contamination the commenters warned about. And routing earns far more on the fresh set: on the old tasks it buys ten points over the cheap model, on the fresh tasks it buys thirty-three, while escalating to the expensive model a fifth of the time instead of never. Routing pays off hardest exactly where the models are tested on work they can't have memorised.

The bit we're not proud of

We caught the harness truncating the models

The first fresh run looked catastrophic, and it was wrong. Our output budget was 4,096 tokens, which is plenty for "write this short function" and nowhere near enough for a hard competitive-programming problem. The reasoning models thought at length, hit the ceiling mid-thought, and emitted no code at all. Every one of the frontier model's fresh failures was a blank page. We only spotted it because we publish every response and could read them. We made the budget configurable, gave the hard suite room, and re-ran. Truncation dropped from every failure to a single task. A closed benchmark makes that same mistake and ships the number, and nobody's the wiser.

How fresh is fresh

The dates behind the tag

Calling a split "fresh" is a judgement, so here's the evidence. The 15 fresh tasks are LiveCodeBench problems published between 2025-01-04 and 2025-04-06, median 2025-02-22, filtered to anything dated 2025-01-01 or later. The model providers don't publish exact training cutoffs, so we hand you the release dates and let you judge how far past the line this run sits, rather than a bare clean-or-contaminated call. Those dates ship in results.json and re-derive under omnisbench verify like everything else. The measured version is the perturbation-gap probe, described step by step below.

The measured probe

How the perturbation-gap test works

  1. Take a fresh problem and its exact test cases, the input-to-output pairs from LiveCodeBench.
  2. A separate rewriting model, never the one being graded on that problem, rewrites the statement. It changes the prose, the variable names and the surface story, and holds the input format, the output format and every constraint word for word. The test cases never change, so the correct answer is identical.
  3. Score every model twice, on the original and on the reworded copy. Grading runs the model's code against the same unchanged tests, so a pass means the same thing both times.
  4. Drop any reworded problem no model can solve, a sign the rewrite broke it, and any problem no model solved at all.
  5. The gap is the drop, original accuracy minus reworded accuracy. A model that leaned on the exact wording falls. A model that understood the problem holds.

The pool spans two providers, and a rewriter can favour or punish its own family's style, so we reword the whole pool twice, once with claude-sonnet-5 and once with gpt-5, and read each model's gap under each. That shows how much the rewriter itself moves the number, not the model.

ModelGap, reworded by claude-sonnet-5 (n=15)Gap, reworded by gpt-5 (n=17)
claude-opus-5+20.0pp+5.9pp
claude-haiku-4-5+40.0pp+5.9pp
gpt-5+20.0pp+23.5pp
gpt-5-nano+20.0pp+52.9pp

One number survives changing the rewriter: gpt-5 gives up about 20 points whoever rewords it, a real sign its score on these recent problems leaned on the phrasing. Every other gap swings with the rewriter. The Anthropic models barely move under gpt-5 and fall hard under sonnet, gpt-5-nano the other way, so at 15 to 17 problems the rewriter and the surviving tasks move the number more than the model does. What holds for all four: every model loses at least 6 points when a fresh problem is reworded, so a post-cutoff date is necessary and not the whole story. Every reworded prompt and response is published, and omnisbench verify re-derives the gaps offline. Both runs and the method: docs/perturbation-probe.md.

Read it honestly

Caveats, up front

Small samples

20 contaminated and 15 fresh tasks. It's a first signal, not a verdict, and the suite grows from here.

Difficulty is tangled with freshness

The fresh set is competitive-programming, which is both newer and harder than grade-school maths. Part of the drop is difficulty, not contamination on its own. We can't fully separate the two yet.

One residual truncation

Even at 16k tokens, one fresh problem still ran the model out of room. That one's genuine difficulty now, not a harness fault.

oracle is a ceiling

It's chosen with hindsight. A shippable router can't be this good. The gap to it is the real scorecard.

Check it yourself

Re-grade the run offline

The whole run, responses included, is in the repo. verify re-runs the graders against the published answers and re-derives this table with no API calls.

# re-grade the corrected fresh run, no keys needed
$ pip install -e .
$ python -m omnisbench.cli verify --run runs/fresh-16k-2026-08-20
VERIFY OK

See the run and the code on GitHub →
Back to the full leaderboard →