All ten best-value models held their positions this week; one combination matched top quality for 97% less cost.
We keep a scoreboard of AI models: how well each does a particular kind of work, and what it costs per thousand requests. This week every model on the scoreboard was re-checked with fresh examples. None had deteriorated. Three newly released models were tested. None beat the models already on the scoreboard. The standout finding: for answering questions from retrieved documents, a combination of measured models delivered quality as high as the best single model for 97% less money.
Week 35 of 2026: ten canaries (small weekly re-checks of four examples each) tested every frontier (the short list of best-value models for a given cluster, meaning a kind of work). Nine frontiers held their positions. One check was inconclusive because no examples were graded. The headline number: for answering questions from retrieved documents, a combination of measured models scored 0.980 quality (an exam score where 1.0 is perfect), matching the best single model within the margin of error (how sure we are, written as give or take 0.004), at 97% lower cost per thousand requests.
This week's frontiers
One row per kind of work. Routed pick is the model Potion currently sends that work to. Stored quality is its exam score when it was measured in full, give or take the margin. Canary is this week's small re-check: a few fresh tasks, scored the same way, to catch a model that has got worse. Held means the re-check landed inside the margin.
| kind of work | verdict | routed pick | stored quality | canary |
|---|---|---|---|---|
| agentic-tool-use | inconclusive | or-ling-3.0-flash | 0.886 ± 0.136 | — n=0 |
| classification | held | name withheld | 0.976 ± 0.033 | 1.000 n=4 |
| code-gen | held | name withheld | 0.979 ± 0.019 | 0.944 n=4 |
| code-review | held | or-deepseek-v4-flash-0731 | 0.869 ± 0.089 | 0.750 n=4 |
| creative | held | or-deepseek | 0.850 ± 0.071 | 0.800 n=4 |
| extraction | held | name withheld | 0.962 ± 0.017 | 1.000 n=4 |
| multi-step-reasoning | held | or-ling-3.0-flash | 0.960 ± 0.055 | 1.000 n=4 |
| rag-answer | held | name withheld | 0.907 ± 0.078 | 0.750 n=4 |
| rewrite-edit | held | or-gpt-mini | 0.866 ± 0.046 | 0.975 n=4 |
| summarization | held | or-ling-3.0-flash | 0.900 ± 0.050 | 0.925 n=4 |
Nine frontiers held their positions. Zero moved to a different model. One check was inconclusive. A canary is a small weekly re-check of four examples that detects whether a model has collapsed, not whether it has shifted by a single percentage point. Every quality figure comes with a margin of error (a 95% confidence interval, meaning we are 95% sure the true score falls within that range). This week or-deepseek-v4-flash-0731 held code-review work at 0.869, give or take 0.089. Or-deepseek held creative writing at 0.850, give or take 0.071. Or-ling-3.0-flash held multi-step reasoning at 0.960, give or take 0.055, and summarization at 0.900, give or take 0.050.
Auditions
An audition is a newly released model's first exam, on the kind of work it looks suited to. It earns a place only by beating the model already doing that work on quality or price.
Three newly released models received auditions (a first exam to see if they beat the current frontier pick). Or-ling-2-6-flash was tested on rewrite-edit work in the multilingual lane and did not beat the incumbent. Or-l3-lunaris-8b and or-nex-n2-mini were both tested on classification work in the small/cheap lane. Neither beat the incumbent.
- or-ling-2-6-flash · multilingual · rewrite-edit · did not beat the incumbent
- or-l3-lunaris-8b · small/cheap · classification · did not beat the incumbent
- or-nex-n2-mini · small/cheap · classification · did not beat the incumbent
Mixing
Mixing means using two or three cheaper models together in a particular way instead of one expensive one. We score combinations by replaying results we already have, so this costs nothing to explore. How a combination works is part of the product and is not published; what it achieves is.
Eight combination findings this week, all from replay (re-scoring a combination from results we already have, without spending money on new requests). For answering questions from retrieved documents, a combination of measured models scored 0.980 over 50 items, 0.004 below the best single model, at 97% lower cost. For code work, a combination scored 1.000 over 90 items, 0.003 above the best single model, more than eight times cheaper. For structured output work, a combination scored 1.000 over 80 items, matching the best single model, between four and eight times cheaper.
- rag-answer · frontier-candidate · quality 0.980 · 97% cheaper · n=50
- code work · cheaper-and-as-good · quality 1.000 · more than 8× cheaper · n=90
- writing work · cheaper-and-as-good · quality 0.936 · between 4× and 8× cheaper · n=14
- structured output work · cheaper-and-as-good · quality 1.000 · between 4× and 8× cheaper · n=80
- agentic-tool-use · frontier-candidate · quality 0.995 · -45% cheaper · n=38
- writing work · frontier-candidate · quality 1.000 · under 1.5× cheaper · n=14
- code work · frontier-candidate · quality 0.995 · under 1.5× cheaper · n=42
- extraction · frontier-candidate · quality 0.987 · -165% cheaper · n=74
What it means for you
If you pay for AI by the request, the gap between the best model and the cheapest model that passes your quality threshold is where your budget drains. This week that gap was measured again across ten kinds of work. It remains wide. Routing each request to the cheapest model that scored above your threshold is how you keep the quality and stop paying for capability you do not need.
How the numbers are made
How these numbers are made, in plain words. We sort requests into kinds of work: sorting text into categories, pulling fields out of documents, writing code, answering from a set of documents, and so on. For each kind we keep a private exam of tasks the models have never seen. Every model sits the same exam. Code is marked by running it; other answers are marked against a reference answer. A model's quality is its average mark, and because an exam is a sample we also give a margin of error: two models whose margins overlap are called a tie. Cost is what a thousand requests would cost at the provider's public prices. A frontier is the short list of models that are the best deal at their level of quality, meaning nothing else is both better and cheaper. Each week we re-check every model on that list with a few fresh tasks to catch any that have got worse, and we give newly released models the exam for the work they look suited to.
- Every quality figure is a mean over a retrieval-hostile suite with a bootstrap 95% interval; two points whose intervals overlap are reported as tied.
- Code-generation and code-review quality is scored by executing the code; other clusters are scored by a rubric against a reference answer.
- A canary is a small weekly sample (four items) against the stored measurement; it detects collapse, not one-point movement.
- Combinations of models are replayed from stored item-level results; agreement between models is modelled conservatively, so the combination figures understate rather than overstate.
- Free-tier listings are excluded from auditions: their quality is not stable enough to measure.
Numbers
10 canaries · 9 held · 0 moved · 1 inconclusive · 36 items graded · 326 listings screened · 3 measured · $1.42 spent
Questions
- What is a routing frontier?
- The short list of model options for one kind of work where no other option beats all of them on quality, cost per thousand requests, and speed at the same time. A routing policy picks from this list: cheapest above a quality floor, fastest, or highest quality.
- Did any model get worse this week?
- No. All ten frontier picks reproduced their stored quality within the 95% margin of error on a four-item canary. One check was inconclusive because no examples were graded for that cluster.
- Which models held their frontier positions this week?
- Or-deepseek-v4-flash-0731 on code-review work at 0.869, give or take 0.089. Or-deepseek on creative writing at 0.850, give or take 0.071. Or-ling-3.0-flash on multi-step reasoning at 0.960, give or take 0.055, and on summarization at 0.900, give or take 0.050. Or-gpt-mini on rewrite-edit work at 0.866, give or take 0.046. Some picks are not named; their numbers are published.
Glossary. A frontier is the short list of models that are the best deal at their level of quality: nothing else is both better and cheaper. A floor is the lowest exam score you are willing to accept. A margin (or interval) is how far the true score could sit from the measured one, because an exam is a sample. A canary is a small weekly re-check. See the docs and the evidence.
This issue was drafted by sending one request to Potion's own API, the same way a customer would. The receipt that came back: kind of work creative, strategy 07b4dc72, policy max_quality, 4970 tokens. We use what we sell.