OmnisBench vs GitHub HydraFusion

A vendor's own numbers, and a benchmark anyone can re-run.

HydraFusion is GitHub's cost-cutting router inside Copilot CLI. By GitHub's own figures it cut cost on all three of its benchmarks and held Opus 5 quality on one of the three. That's a real result, and the fact that Microsoft shipped it tells you the category matters. OmnisBench does the other half of the job: it measures how much routing efficiency is genuinely there, on fresh tasks the models can't have memorised, with every number re-gradable offline. Apache-2.0.

At a glance

How they compare

DimensionHydraFusion (GitHub Copilot)OmnisBench
TypeCost-cutting routing feature inside Copilot CLIOpen-source benchmark you run yourself
Numbers come fromGitHub, on GitHub's own benchmarksAnyone, re-graded offline from published responses
Works withGitHub Copilot onlyAny model pool you configure, tool-agnostic
TasksTerminalBench 2.1, DeepSWE, CheckpointBenchFresh, contamination-controlled tasks
Independently verifiablePublished result, not re-runnable by youYes, omnisbench verify re-grades offline, zero API calls
LicenseCopilot product featureApache-2.0
The published result

Cheaper everywhere, quality held in one of three

Here are GitHub's own numbers, and they're honest about the trade-off. Cost fell on all three benchmarks. Opus 5 quality held on one.

BenchmarkCost vs Opus 5Quality vs Opus 5
TerminalBench 2.167% lowerup 4.9 points
DeepSWE36% lowerdown 1.5 points
CheckpointBench65% lowerdown 0.1 points

Cheaper is the easy half to prove. The hard half is what it costs you in quality, and on two of three it cost a little. That's the number a benchmark should pin down on tasks a model hasn't already seen, which is exactly what OmnisBench is for.

Whose numbers

A vendor grading its own homework, and a check anyone can run

GitHub ran HydraFusion against its own chosen benchmarks and reported the result. That's normal, and the figures look reasonable. But a router's whole value is a cost-quality trade-off, and the honest way to know where that lands is a test the vendor doesn't own. OmnisBench publishes every model response for every task in results.json, and omnisbench verify re-runs the graders against those responses and re-derives the leaderboard offline, with zero API calls. Change one stored answer and it fails. You can check a number yourself in minutes rather than take it on trust.

One tool, or all of them

HydraFusion stops at the Copilot boundary

HydraFusion only works inside GitHub Copilot CLI, billed per token on a Copilot plan. If your developers use Claude Code, Cursor, Codex, or hit the Anthropic and OpenAI APIs directly, it can't see them. OmnisBench measures routing efficiency independently of any one tool. OmnisRouter then acts on it across all of them, with a real local spend cap and a kill switch, and OmnisVigil turns the whole fleet's spend into a number a finance team can read. HydraFusion has no budgets, caps or per-team attribution. It lowers a request's price. It doesn't tell you who spent what.

Who each is for

Pick based on the job

HydraFusion

Your team lives entirely inside GitHub Copilot, you want a lower per-request cost with no setup, and you're happy to take GitHub's word on the trade-off.

OmnisBench

You want to know how much routing efficiency is actually available, on fresh tasks, with a number you can re-grade yourself, whatever tool the routing lives in.

See the numbers

OmnisBench's fresh split

On tasks the models can't have memorised, ideal routing hit 93.3% task success, above the frontier model itself, at about 60% lower cost, and every figure re-grades offline. We're putting HydraFusion through the same open test and will publish where it lands.

See the full leaderboard on the homepage →
Reproduce it on GitHub →
Read the open-source LLM routing benchmark overview →

Questions

Frequently asked

Is OmnisBench an alternative to HydraFusion?
They do different jobs. HydraFusion is a cost-cutting routing feature inside GitHub Copilot CLI. OmnisBench is an open benchmark that measures how much routing efficiency is available across today's models, on fresh tasks, with every number re-gradable offline. Use OmnisBench to check a cost-quality claim like HydraFusion's rather than take it on faith. For a router that works across Claude Code, Cursor, Codex and direct API, see OmnisRouter; for team spend governance, see OmnisVigil.
Does HydraFusion cut costs without hurting quality?
By GitHub's own published numbers, HydraFusion cut cost on all three of its benchmarks (67% on TerminalBench 2.1, 36% on DeepSWE, 65% on CheckpointBench) but held Opus 5 quality on one of the three (up 4.9 points on TerminalBench, down 1.5 on DeepSWE, down 0.1 on CheckpointBench). So on two of three it was cheaper and slightly worse. Those are GitHub's benchmarks, run by GitHub. OmnisBench is the open way to see where a router lands on tasks the models can't have memorised.
Does HydraFusion work outside GitHub Copilot?
No. HydraFusion runs inside GitHub Copilot CLI on Copilot plans. It does nothing for Claude Code, Cursor, Codex, or direct Anthropic and OpenAI API spend. OmnisBench measures routing efficiency independently of any one tool, and OmnisRouter routes across all of them.