Code generation: a model at 1/270th the price of the best scorer, at 99% of its quality.
Research · weekly

Frontier Notes

Weekly measurements of which AI models are cheapest at a given quality, and what routing between them saves.

Every issue is produced from the same measurements that route production traffic: a weekly drift check on every routing frontier, auditions of newly listed models, and replayed combinations of measured models. Numbers carry intervals; names that are part of the product are withheld, their numbers are not.

2026-W34 · 2026-08-22

All ten best-value models held their positions this week; combining cheaper models matched top quality on code work for a ninth of the cost.

Every week we test whether the best-value AI model for each kind of work is still the best value. This week all ten held steady. The finding that matters: on code work, using cheaper models together matched the most expensive option for less than a ninth of its price.

10 canaries · 10 held · 0 moved · 3 new models measured · 9 combination findings

RSS: /research/feed.xml · Method: every issue carries the same method note; see the docs for the taxonomy and scoring.

Frontier Notes — weekly measured model routing research · Potion