Discover why enterprises prefer cheaper AI models over frontier ones. Latest data shows 6 models deliver 77% performance at 2.5% cost of Claude Fable 5.
Why 84% of AI Tokens Aren't State-of-the-Art (2026)
Key Insights
- State-of-the-art models improved two-thirds in capability since November 2025, with new models shipping every three days
- 84% of tokens on OpenRouter bypass frontier models despite rapid SOTA advancement
- Six popular models deliver 77% of frontier performance for only 2.5% of Claude Fable 5's cost ($0.50/m vs $20/m)
- Price elasticity drives adoption: Claude Fable 5 captured just 6% of Anthropic tokens despite strong performance
- The gap keeps closing: best open-weight models reached 80% of frontier score by May 2026, up from 48% a year earlier
The Pace of AI Model Improvement (But Limited Adoption)
State-of-the-art models are two-thirds smarter than they were last November. Labs ship two new models every three days, sustaining a frenetic pace of improvement.
Yet adoption tells a different story. Despite this rapid advancement, 84% of tokens on OpenRouter aren't state-of-the-art. The frontier keeps advancing, but most users aren't following it there.
The Six-Model Supermajority: Performance at a Fraction of the Cost
Six models carry the supermajority of tokens on OpenRouter and deliver about 77% of frontier performance at dramatically lower cost. Their blended price is $0.50 per million tokens, compared to Claude Fable 5 at $20 per million tokens—a 40x difference.
This price-to-performance split reflects how enterprises actually allocate spend. Claude Fable 5 generated roughly 75% as much model-attributed revenue as OpenAI's GPT-5.6 Sol in July, despite being substantially more expensive.
Price Elasticity: Why Users Choose "Good Enough"
Ramp's data reveals strong price elasticity in the market. Fable 5 captured only 6% of Anthropic tokens and 11% of Anthropic spend in its first month after launch. GPT-5.6 Sol, OpenAI's priciest mainline tier, held about a quarter of OpenAI tokens.
Performance is already good enough at a meaningful discount. The best open-weight models reached 80% of frontier score by May 2026, up from 48% a year earlier. As the performance gap closes from below, users and enterprises optimize against a different Pareto frontier—price over performance rather than maximum capability.
Application Deployment: The Real Frontier
Frontier models still win on software architecture and security design, where the best available model earns its premium price. But in application deployment, the calculus shifts.
More portfolio companies and startups default to smaller models, fine-tuned models, and open source. They optimize against price over performance because good enough truly is good enough for their use cases.
If share stops shifting and good enough stays good enough, the economics of state-of-the-art training runs face pressure. A nine-figure training investment must win market share to pay for itself, and that bar rises with time.
Conclusion
The AI frontier moves faster than ever—models improve, labs ship constantly, and capability gaps shrink. But the market's actual behavior suggests a different story: enterprises vote with their tokens for models that deliver meaningful performance at a fraction of frontier cost. The real Pareto frontier isn't between best and rest—it's between price and performance, and for most applications, cost wins.
Original source: Honestly, Who Buys SOTA?
powered by osmu.app