Most AI tokens use cheaper models delivering 77% of frontier performance at 2.5% of the cost. Here's why "good enough" is winning in 2026.
Why 84% of AI Tokens Aren't Using State-of-the-Art Models
Key Insights
- State-of-the-art models are 67% smarter than November 2025, with major-lab releases averaging 20 per month
- 84% of tokens on OpenRouter use non-frontier models that deliver 77% of frontier performance at just 2.5% of the cost
- Six popular models carry 80% of volume with a blended price of $0.50 per million tokens versus Claude Fable 5's $20
- Price elasticity is real: Fable 5 captured only 6% of Anthropic's token volume despite 11% of spending within a month of launch
- The gap is closing: Best open-weight models reached 80% of frontier performance by May 2026, up from 48% a year earlier
The Frontier Keeps Climbing
State-of-the-art models have improved dramatically. The pace of advancement remains relentless, with two major new models shipping roughly every three days. Performance gains—measured on the Artificial Analysis Intelligence Index—jump 3–5 points every quarter, with smaller steps filling the gaps between major releases.
The "Good Enough" Paradox
Despite rapid frontier progress, the data tells a different story on OpenRouter. Most token consumption concentrates on models that sit well behind the cutting edge. The six models users choose most often generate about 77% of frontier performance but cost just 2.5% of what Claude Fable 5 commands. This isn't a niche behavior—it represents the supermajority of actual token usage.
The blended price of these popular six models sits at roughly $0.50 per million tokens, while Fable 5 costs $20—a 40-fold difference. Performance losses are real but modest enough that the price advantage overwhelms them for most workloads.
Price Elasticity Shapes Market Share
Pricing power has limits. When Anthropic launched Fable 5, it claimed only 6% of Anthropic token volume and 11% of Anthropic spending a month later. By contrast, OpenAI's most expensive model, GPT-5.6 Sol, held roughly 25% of OpenAI tokens. Yet Fable 5 generated about 75% as much model-attributed revenue as GPT-5.6 Sol in July 2026—a sign that cheaper alternatives are capturing meaningful workload share.
The implication is stark: buyers are price-elastic. Each new frontier release should capture less share than the one before it as the performance gap from cheaper options shrinks.
Where Frontier Models Still Win
State-of-the-art models aren't irrelevant. They earn their premium in software engineering architecture and security design—domains where the best available option justifies its cost. First-party deployment by OpenAI, Anthropic, and Google also skews toward frontier models on native APIs, a segment OpenRouter data doesn't fully capture.
For application deployment and general use, however, the economics have shifted. More startups and portfolio companies now default to smaller models, fine-tuned variants, and open-source options. They optimize for a different Pareto frontier: price over performance, not performance at any cost.
The Economics of Training Giant Models
A nine-figure training run must capture share to justify its cost. As cheaper models close the performance gap, that bar rises. If market share stabilizes and "good enough" remains sufficient, the economics of frontier models change fundamentally. The 80% performance level from open-weight models—up from 48% a year earlier—signals the frontier is still being chased, but the incentive to pay for it keeps shrinking.
Conclusion
State-of-the-art models are genuinely smarter than they were months ago, and that progress matters. But the real market story isn't at the frontier. It's in the 84% of tokens using models that cost 40 times less while delivering most of the performance. For most builders, frontier models have already become optional—and that economics won't reverse.
Original source: Honestly, Who Buys SOTA?
powered by osmu.app