Explore why GPT-OSS-120b captures 36% of Claude Opus traffic & how mixture-of-experts architecture is reshaping the AI inference market across cost, speed & ...
AI Model Market Segmentation: Why Older Models Still Dominate (2026)
Key Insights
- A one-year-old open-source model (GPT-OSS-120b) serves 36% of Claude Opus 4.8's daily token volume on OpenRouter, proving the inference market has fragmented across cost, speed, and accuracy
- Market segmentation is accelerating with Anthropic's Opus 5 launch targeting mid-market competition and Poolside's Laguna S 2.1 bringing frontier-class quality to local deployments
- Mixture-of-experts (MoE) architecture enables 118B-parameter models to run at 26B model speeds, collapsing the traditional boundary between local and frontier tiers
- Real-world accuracy improvements matter: Laguna S 2.1 reduced tool-call failure rates from 27.1% to 20.1% compared to smaller models, demonstrating practical competitive advantage
- OpenRouter token data reveals surprising buying patterns: GLM 5.2 leads at 495B daily tokens, with older open-weight models capturing substantial share alongside frontier models
The Token Market Has Fragmented
OpenRouter's token serving data tells a striking story: GLM 5.2 tops the field at 495B tokens daily, while Claude Opus 4.8 serves 199.6B, and the year-old GPT-OSS-120b holds 71.3B tokens—roughly one-third of Opus traffic. This isn't a market failure; it's evidence of healthy segmentation.
Buyers are making deliberate tradeoffs. Some need frontier accuracy for complex reasoning. Others prioritize cost-per-token. Still others optimize for speed or local deployment. The market now offers multiple tiers: small models (26B parameters), mid-range dense models, sparse mixture-of-experts architectures, and frontier-class systems.
Why Older Models Still Win Market Share
A model released in August 2025 competing successfully against Opus 4.8 (shipped weeks ago) reflects deeper market realities. Different use cases require different models. OpenRouter's serving patterns show:
- Cost-conscious users gravitate toward older open-source models despite newer alternatives
- Specialized tasks (coding, vision, multimodal) drive selection more than release date
- Origin matters: US-built models compete against Chinese models (Kimi 3, GLM 5.2) in specific regions
- Architecture choice (dense vs. sparse) determines deployment environment—local inference on M5 Max favors mixture-of-experts
Mixture-of-Experts Collapses the Speed-Accuracy Tradeoff
The competitive shift accelerated when Laguna S 2.1—a 118-billion-parameter mixture-of-experts model—runs at the same decode speed as a 26-billion-parameter dense model. This happens because only 8 billion parameters activate per token, despite 118 billion parameters in memory.
In production testing, this translates to measurable improvement: tool-call failure rates dropped from 27.1% (Gemma 4 26B) to 20.1% (Laguna S 2.1)—a 7-percentage-point reduction in error rates for a coding and email agent driving real automation workflows.
That boundary—where frontier-class accuracy now runs locally—represents a permanent shift. Mixture-of-experts lets smaller hardware run larger models, expanding the addressable market for both open-weight and commercial models.
The Competitive Acceleration
Recent launches demonstrate this intensifying segmentation:
- Anthropic shipped Opus 5 explicitly smaller and cheaper than its predecessor, targeting mid-market buyers and directly contending with Moonshot's Kimi 3
- Poolside released Laguna S 2.1, a US-built mid-market alternative with strong local inference properties
- Frontier models no longer need to serve all token types—they concentrate on genuinely hard problems while specialized models handle the rest
Conclusion
OpenRouter's traffic patterns reveal a maturing market: segmentation across cost, speed, accuracy, and origin is the sign of healthy competition, not fragmentation. The year-old GPT-OSS-120b persists because it's genuinely fit-for-purpose for many buyers. Mixture-of-experts architectures enable local deployment with frontier-class quality. And Anthropic's latest launch confirms competitive pressure is pushing the entire market forward.
The AI inference market is no longer a single ladder—it's a marketplace with multiple aisles, each optimized for different needs.
Original source: Yeltsin in the AI Aisle
powered by osmu.app