The right model(s) for every request.
six real decisions from the measured frontier — not a simulation · click a row for its receipt
No. A gateway answers how do I call any model? Potion answers which model does this request deserve? Those are different layers, and the second one is where the money is.
Access stopped being scarce the day gateways shipped. Judgment — measured per kind of work, stated with its error bars, enforced as a floor — is the scarce layer. That layer is Potion, and it works the same over any gateway or provider underneath.
Five models writing code to specification — scored by running their code, not by opinion. The quality difference across this table is two points in a hundred. The price difference is two hundred and seventy fold.
This is why routing pays: most requests deserve the bottom row, a few genuinely need the top one, and only a measurement can tell them apart.
measured 2026-08-20 · retrieval-hostile suite · scored by execution · error bars on the full table in the docs
↑ the routed pick — 99% of the top row's quality at 1/270th the price. The name? That's the product.
bars show cost per 1,000 requests · teal = what a 0.95 quality floor actually buys
A cheap model answers and reports how sure it is. Only when that confidence falls below a measured threshold does the request go on to a stronger one. Most traffic never reaches the expensive model at all — so the combination can land at the top of the measured quality range while costing a fraction of sending everything to the strong model.
There is nothing for you to assemble. You set the same rule you would set anyway — stay above this quality, stay under this cost — and if a combination is the best way to honour it, that is what serves your request. The receipt names whatever answered.
There are far more useful combinations than there are models, and almost none of them have been measured by anyone. That is the dimension this company is named for.
Measured options for one kind of work, error bars included. Pick a rule, drag the slider, and you are running the same selection the router runs in production — including its refusal to answer when nothing measured qualifies.
or-deepseek-v4-flash-0731 — quality 0.98 ±0.04, $0.0304 per 1k requests, p95 9,702 ms.
Real measured points, quoted from the committed frontier — hover any dot for its name and numbers. The cheapest row costs under a cent per 1k and measures 0.50, a coin flip, which is exactly why the router will not send reasoning work there: cheap only wins where the measurement clears your floor. Try latency_bound at 2,500 ms — the answer changes.
Usage-based today. The direction of travel is to charge against measured savings — the only pricing that stays honest when the whole point of the product is spending less, and the only one that makes the bill fall when we do our job badly.
Customers bring no accounts and no keys. Potion buys from every provider at once, which is also what lets it reach the whole market rather than the one account a customer happened to open.
Point your client at Potion. Nothing else about your app changes.
Each request goes where the evidence says it should, not where habit does.
A line their traffic is never allowed to fall below, enforced per request.
What was chosen, why, and which measurements it relied on.
The request is refused before the money is spent, never after.
Spend is attributable to decisions you can audit, line by line.
Point a client at Potion and watch the routing decisions arrive with the answers.