This week we measured a model at 1/270th the price of the best — at 99% of its quality. The name is the product.
Measured model routing

Cut your AI bill in half.

The right model(s) for every request.

  • AI bills fall 49% when Potion routes each request to the cheapest model measured good enough.
  • Cost, quality, or speed: you set the rule, Potion picks from the measured Pareto frontier.
  • Your routing stays current automatically: new models are measured on release, and our research finds model combinations nobody else has.
  • One line of code. Every answer carries a receipt: what ran, and why.
receipt
kind of work
classification
routed to
or-deepseek-v4-flash-0731
price · quality
$0.0201/1k · q 0.99 94%
why
cheapest point within a point of the perfect scorers
routing 6 measured requests…
requestrouted toprice/1kquality

six real decisions from the measured frontier — not a simulation · click a row for its receipt

The obvious question

Isn't this what OpenRouter does?

No. A gateway answers how do I call any model? Potion answers which model does this request deserve? Those are different layers, and the second one is where the money is.

You get
gateway · Every model, one API
The right model for each request
Who chooses
gateway · You do, once per app
The measurements do, per request
Based on
gateway · Leaderboards, habit, vibes
Held-out tests of your kind of work
Quality
gateway · Whatever you picked
A floor your traffic never falls below
After the answer
gateway · Tokens and a price
A receipt: what served it, and why
A new model ships
gateway · You re-evaluate by hand
Measured first, adopted only if it earns it

Access stopped being scarce the day gateways shipped. Judgment — measured per kind of work, stated with its error bars, enforced as a floor — is the scarce layer. That layer is Potion, and it works the same over any gateway or provider underneath.

01
Held-out sets
Graded on items the models never see in advance.
02
Known answers
Scored against a reference, or by running the code.
03
Stated uncertainty
Every score carries the interval its evidence supports.
04
Re-measured on release
A new model is tested before it is ever routed to.
Measured, not claimed

The same work.
A 270× price range.

Five models writing code to specification — scored by running their code, not by opinion. The quality difference across this table is two points in a hundred. The price difference is two hundred and seventy fold.

This is why routing pays: most requests deserve the bottom row, a few genuinely need the top one, and only a measurement can tell them apart.

measured 2026-08-20 · retrieval-hostile suite · scored by execution · error bars on the full table in the docs

or-grok-4.6xAIquality 1.000 · $6.2568/1k
or-gemini-flashGooglequality 0.996 · $0.5506/1k
or-gpt-miniOpenAIquality 0.990 · $0.2509/1k
or-deepseekDeepSeekquality 0.980 · $0.1904/1k
or-████████████name withheldquality 0.979 · $0.0231/1k

↑ the routed pick — 99% of the top row's quality at 1/270th the price. The name? That's the product.

bars show cost per 1,000 requests · teal = what a 0.95 quality floor actually buys

270×

the price range across five models doing the same measured work — quality within two points in a hundred

99%

of the most expensive model's quality, from the pick our measurements route to, at 1/270th the price

49%

the measured saving across a typical traffic mix, with a quality floor enforced on every request

$0.0231

per 1,000 requests — the routed pick, name withheld

$6.2568

per 1,000 requests — the model most teams would have picked

same suite · scored by execution

Both rows did the same held-out work and were scored the same way — by running what they wrote. The gap between them is not quality. It is the cost of choosing without measuring.

The problem

The same job can cost 170 times as much, for no benefit.

There are hundreds of AI models. They differ enormously in price and only sometimes in quality — and which one is best changes every few weeks as new ones ship.

Almost nobody re-checks. A team picks a model once, wires it in, and keeps paying that price on every request forever — including the thousands of easy ones a model costing a fraction as much would answer just as well.

for engineers
The chart is the measured extraction frontier: quality is a graded score over 74 held-out items, cost is USD per 1,000 requests at each strategy's measured token profile.
one task · pulling data out of documentsevery bar does the job
$0.02
q 0.960
$0.20
q 0.961
$0.29
q 0.956
$0.37
q 0.960
$0.58
q 0.984
$1.44
q 0.961
$1.46
q 0.973
$1.58
q 0.978
$1.77
q 0.984
$4.07
q 0.978

Every bar does the job. The dearest costs 172× more than the cheapest and measures 1.8 points better out of 100. Most companies are on a bar near the bottom of this chart for every request they send, because they picked one model and moved on.

What it is worth

Set how good the answers have to be. See what routing saves.

This is not a marketing calculator. It runs Potion's real selection rule over Potion's real measurements, live, as you drag. Raise the quality bar and watch categories drop out — including the ones we would refuse to take money for.

savings model · runs the real selection rule, live

$50k a month is $600k a year.

0.80

Measured quality, 0 to 1 — the floor Potion is never allowed to go below.

Cut from your AI bill
84%
versus running the best model on everything
Saved per year
$501k
$600k a year becomes $99k
Where the saving comes from
Writing code
100%
Sorting and tagging
99%
Pulling data out of documents
96%
Rewriting and editing
92%
Multi-step reasoning
90%
Agents using tools
86%
Summarising
80%
Answering from your documents
78%
Creative writing
47%
Reviewing code
45%

Some rows sit at or near zero and that is the model working. On multi-step reasoning the cheap options measure 0.46 against 0.98 for the expensive one — so Potion pays for the expensive one and saves you nothing there. Half the categories carry most of the saving; being told which half is the product.

Computed live from Potion's own measurements, against a baseline of running the highest-quality model on everything. It weights every kind of work equally — a real bill depends on your traffic mix, which is the first thing Potion measures once you connect.

Why this is hard to copy

Anyone can call a cheaper model. Knowing when that is safe is the asset.

The router is a week of engineering. The evidence it routes on is not — and it is the half that compounds.

the corpus accumulates

The measurements compound

Every model, on every kind of work, graded against known answers — and re-graded when a new one ships. That corpus is the product, and it is worth more every week than it was the week before.

Content-addressed: re-measuring an unchanged pair costs nothing, so the corpus only ever pays for what is genuinely new.
σ
uncertainty is carried, not dropped

Benchmarks are the wrong instrument

Public leaderboards rank models on average, across work you do not do. Potion measures each model on each kind of job, and keeps the error bars — the only comparison that can tell you what to send where.

A separate quality/cost/latency frontier per kind of work. Selection is always per-workload, never global.
it improves in place

It gets better the longer you run it

You start on measurements of your kind of work. As your own traffic accumulates, the routing retunes to it specifically — so the saving grows without you changing a line or paying more attention.

Shared measurements serve from day one; measurements of your own traffic take over as they accrue.
How it works

Read the request. Pick the cheapest model good enough. Show your working.

01

It reads what you are asking for

Before anything is chosen, the request is sorted into a kind of work — writing code, summarising, pulling data out of a document, and so on. Each kind has its own answer about which model is best.

cluster = code-gen
02

You set the rule, it picks the model

You say what matters: never go below this quality, or never spend above this much, or never take longer than this. Potion holds a measured map of every option — single models and combinations of them — and picks the best one that obeys your rule.

policy = min_cost · qualityFloor 0.80
03

Every answer comes with a receipt

The response carries what it chose, why, and which measurements it relied on. If Potion had nothing measured for your request, the receipt says that too rather than quietly guessing.

x-frontier-trace: cluster=code-gen;strategy=6efe8a56;frontier=v2;policy=min_cost;fallback=0
Sometimes the answer is not one model

A mixture can beat anything you could have picked.

A cheap model answers and reports how sure it is. Only when that confidence falls below a measured threshold does the request go on to a stronger one. Most traffic never reaches the expensive model at all — so the combination can land at the top of the measured quality range while costing a fraction of sending everything to the strong model.

Potion measures those combinations exactly the way it measures single models: the same held-out items, the same known answers, the same error bars. A mixture earns a place on the map or it does not appear on it.

There is nothing for you to assemble. You set the same rule you would set anyway — stay above this quality, stay under this cost — and if a combination is the best way to honour it, that is what serves your request. The receipt names whatever answered.

There are far more useful combinations than there are models, and almost none of them have been measured by anyone. That is the dimension this company is named for.

for engineers
Cascade, ensemble, draft-verify, best-of-n and staged-upgrade composites are all first-class strategies with their own executors, generated against the model registry and promoted only on a paired bootstrap over held-out items — the same gate a single model faces.
a cascade · measured as one strategy
requesta cheap modelanswers first?sure enough?yesdonenoa stronger modelonly when neededdone
How you know we are not making this up

It tells you when it cannot tell.

Every quality score Potion reports comes with a margin of error, the way a poll does. When two models are close enough that the measurement cannot separate them, Potion says they are tied instead of inventing a winner — and then picks the cheaper one.

two real points · error bars overlapping · reported as a tie
0.800.850.900.951.00or-sonnetor-gpt-minishared ground

Too close to call. Those two bars overlap almost completely, so the quality difference is not real evidence — while the cheaper one costs 46% less and answers 1.9× faster. A leaderboard would rank them and let you overpay for the gap.

for engineers
Two real points on the rewrite-edit frontier: or-sonnet at 0.921 ±0.062 against or-gpt-mini at 0.879 ±0.072. The rule underneath is that no dial is exposed without a measurement beneath it — an unmeasured model stays reachable and is never auto-selected.

The actual map, if you want to drive it yourself

Measured options for one kind of work, spanning a hundredfold price range. Pick a rule, drag the slider, and you are running the same selection the router runs in production — including its refusal to answer when nothing measured qualifies.

the measured map · one kind of work · selection runs as you drag
0.40.60.81.0$0.01$0.1$1cost per 1,000 requests (log)measured qualityfloor 0.60or-████████████ — q 0.50 ±0.14 · $0.0093/1k · p95 7,160 msor-████████████or-deepseek-v4-flash-0731 — q 0.98 ±0.04 · $0.0304/1k · p95 9,702 msor-deepseek-v4-flash-0731or-nemotron-3.5-lightning — q 0.76 ±0.12 · $0.0845/1k · p95 2,739 msor-nemotron-3.5-lightningor-gpt-full — q 0.56 ±0.14 · $0.1482/1k · p95 1,126 msor-gpt-fullor-gemini-flash — q 0.66 ±0.13 · $0.1656/1k · p95 2,293 msor-gemini-flashor-inkling-small — q 0.92 ±0.08 · $0.1906/1k · p95 4,009 msor-inkling-smallor-gemini-3.7-flash — q 1.00 ±0.00 · $0.2929/1k · p95 3,242 msor-gemini-3.7-flashor-inkling — q 0.94 ±0.07 · $0.5657/1k · p95 2,449 msor-inklingor-kimi-k3 — q 0.96 ±0.05 · $1.65/1k · p95 2,422 msor-kimi-k3

or-deepseek-v4-flash-0731 — quality 0.98 ±0.04, $0.0304 per 1k requests, p95 9,702 ms.

x-frontier-trace: cluster=multi-step-reasoning;strategy=1fd419ee;frontier=v3;policy=min_cost;fallback=0;provenance=live

Real measured points, quoted from the committed frontier — hover any dot for its name and numbers. The cheapest row costs under a cent per 1k and measures 0.50, a coin flip, which is exactly why the router will not send reasoning work there: cheap only wins where the measurement clears your floor. Try latency_bound at 2,500 ms — the answer changes.

The business

We get paid out of what we save you.

Usage-based today. The direction of travel is to charge against measured savings — the only pricing that stays honest when the whole point of the product is spending less, and the only one that makes the bill fall when we do our job badly.

Customers bring no accounts and no keys. Potion buys from every provider at once, which is also what lets it reach the whole market rather than the one account a customer happened to open.

what a customer gets on day one
One line of code

Point your client at Potion. Nothing else about your app changes.

Routing on every measured kind of work

Each request goes where the evidence says it should, not where habit does.

A quality floor

A line their traffic is never allowed to fall below, enforced per request.

A receipt on every answer

What was chosen, why, and which measurements it relied on.

A spending cap

The request is refused before the money is spent, never after.

A bill that argues for itself

Spend is attributable to decisions you can audit, line by line.

Change one line. Keep the receipts.

Point a client at Potion and watch the routing decisions arrive with the answers.