Defining the AI Model Price Collapse
The AI model price collapse is the rapid drop in per‑token costs across text and image models, driven by open‑weight systems, aggressive API discounts, and a shift in developer demand toward cheaper, high‑volume workloads without major quality trade‑offs. This collapse is visible in both production usage and formal pricing tables. On the usage side, OpenRouter’s monthly leaderboard tracks token consumption across thousands of apps, revealing which models developers rely on when real money and latency are on the line. On the pricing side, Artificial Analysis blends cache hits, input, and output tokens into a single cost figure, turning a messy market into a clean token cost comparison. Together, these datasets show that price and performance are no longer on opposite ends of a trade‑off; instead, they describe a crowded middle where the cheapest AI models still meet enterprise expectations.
OpenRouter Leaderboard: Usage as a Market Signal
OpenRouter’s leaderboard has become a key lens on AI model pricing and adoption because it measures token consumption, not marketing hype. DeepSeek V4 Flash tops the chart with 10.9 trillion tokens and nearly 10x growth, while Tencent’s Hy3 Preview follows closely at 10.7 trillion tokens after going from “zero to near‑parity” in a single month. Claude Opus 4.7 and Claude Sonnet 4.6 occupy strong middle-tier positions, displaying steady enterprise use rather than launch spikes. OpenRouter’s own Owl Alpha also appears high on the list, likely as a default routing target. Overall, text models display a diverse competitive field defined by cost and capability combinations, while image generation sees more consolidation, with Google models dominating traffic. For developers, the leaderboard functions as a live market map, showing where traffic flows when token costs and production constraints collide.
Cheapest AI Models and the New Cost Baseline
Artificial Analysis’ pricing data shows how far the AI pricing war has pushed down the cost of capable models. DeepSeek V4 Flash (Max) sets the tone at a blended USD 0.06 (approx. RM0.28) per million tokens, built on a 284‑billion‑parameter Mixture‑of‑Experts design with only 13 billion parameters active per token. GPT‑OSS‑20B (High) follows at USD 0.07 (approx. RM0.32) per million, trading some reasoning power for extreme throughput on simpler tasks. Higher‑capacity options remain aggressively priced: DeepSeek V4 Pro (Max) and MiMo‑V2.5‑Pro both sit at USD 0.18 (approx. RM0.83) per million tokens, with GPT‑OSS‑120B (High) at USD 0.20 (approx. RM0.92). According to Artificial Analysis, this wave of open‑weight models and budget tiers has “collapsed token costs by an order of magnitude in under two years,” setting a new baseline for everyday workloads.
Token Routing: Cost Efficiency Without Major Quality Loss
Token consumption data shows developers shifting traffic toward these cheaper AI models while maintaining practical quality levels for most workloads. DeepSeek V4 Flash’s 995 percent growth on OpenRouter reflects aggressive adoption for output‑heavy pipelines where speed and volume matter more than perfect factual accuracy. Hy3 Preview’s surge, driven by a USD 0.063 (approx. RM0.29) per million input token price and strong long‑context behavior, hints at its role in agent workflows and complex browsing tasks. At the same time, mid‑priced but higher‑accuracy models like Claude Opus 4.7 and Claude Sonnet 4.6 still command billions of tokens, suggesting developers are mixing tiers: routing routine calls to ultra‑cheap engines while reserving premium models for high‑risk or critical reasoning steps. This blended strategy turns the pricing war into a practical tool for cutting costs without tearing down service quality.
Pricing Transparency, Blended Costs, and Market Consolidation
Pricing transparency tools from Artificial Analysis help developers interpret the AI pricing war in more than headline numbers. Their blended price assumes a 7:2:1 ratio of cache hits, input tokens, and output tokens, which closer matches real API traffic than raw list prices. That framing explains why Mixture‑of‑Experts models with large contexts, like DeepSeek V4 Flash and V4 Pro, can look extremely cheap in practice. In parallel, OpenRouter’s leaderboard hints at consolidation patterns: Google dominates image generation, while text workloads split across open‑weight Chinese models, Anthropic’s Claude family, OpenAI’s OSS series, and platform defaults like Owl Alpha. For enterprises, the combination of detailed token cost comparison and live usage rankings reduces guesswork. Teams can align cache strategies, input truncation, and routing policies with clear evidence of which cheapest AI models deliver the best cost‑per‑useful‑token in open production.






