The New Reality: AI Power Is Cheapening—and That Changes Everything
LLM pricing competition is the emerging dynamic where leading AI providers increasingly compete not on maximum model capability alone, but on lowering token costs, improving AI cost efficiency, and reshaping enterprise AI budgets as organizations deploy generative AI at scale across workflows and applications. This is no minor adjustment; it is a structural change in how the generative AI market will be valued. The headline moment is OpenAI’s decision to sharply cut prices on its latest GPT-5.6 models, turning what used to be a premium capability into something closer to a commodity. That move forces every enterprise AI leader to admit an uncomfortable truth: they may have been overpaying for intelligence they did not fully monetize—and the era of “tokenmaxxing” with fuzzy ROI is ending.

OpenAI’s Aggressive Cuts: A Signal, Not a Discount
OpenAI has significantly reduced the prices of two GPT-5.6 models, escalating LLM pricing competition and intensifying the generative AI market battle over unit economics. Chief Executive Officer Sam Altman said the goal is to offer “the best price/intelligence tradeoff at every level,” a clear admission that price is now strategic, not tactical. GPT-5.6 Luna was cut by 80%, to USD 0.20 (approx. RM0.92) per million input tokens and USD 1.20 (approx. RM5.52) per million output tokens, while GPT-5.6 Terra dropped 20% to USD 2 (approx. RM9.20) and USD 12 (approx. RM55.20) respectively. These reductions apply not only to API calls but also to how usage is calculated for paid Codex and ChatGPT Work subscribers, making advanced AI more affordable for everyday software engineering and automation tasks. Crucially, OpenAI insists this comes from efficiency gains across its stack, not from weakening its models.
A $120 Billion LLM Economy Under Pressure to Pay for Its GPUs
The LLM economy has grown into an estimated USD 120 billion (approx. RM552 billion) annualized market, transforming AI from a research frontier into a commercial engine spanning enterprise software, developer tools and consumer apps used by hundreds of millions of people. According to economist Callum Williams, that revenue figure highlights both the speed of growth and the scale of expectations now placed on generative AI. Yet the number is small compared with the trillions being poured into AI infrastructure, from GPUs and data centers to electricity and networking. Investor pressure is clear: this industry must prove sustainable unit economics, not just dazzling demos. Anthropic currently commands the largest share of LLM revenue, thanks to a focus on business customers and recurring, usage-based income across coding, reasoning and safer enterprise deployments. Meanwhile, open-source strategies, including Llama-style open-weight models, keep narrowing the gap, driving prices down and forcing proprietary vendors to justify every marginal dollar of spend.
From “Most Capable” to “Best Total Cost of Ownership”
For roughly two years, AI competition has been framed as a race to release more capable frontier models; that race is now shifting decisively toward AI cost efficiency and accessibility. Every query runs on expensive GPUs, power and network capacity, so lowering inference cost is no longer a nice-to-have—it is the business model. OpenAI’s explanation is telling: efficiency gains come from better architecture, optimized inference systems, improved routing and tighter context management that cuts unnecessary token generation while keeping performance. These changes reflect a broader architectural trend: separating reasoning engines from memory and tools, letting agents do more work for fewer tokens. The result is a market where intelligence is becoming table stakes, while total cost of ownership becomes the differentiator for enterprise AI deployments. In plain terms, model strength still matters, but CFOs now care far more about cost per million tokens than benchmark scores.
What Enterprises Must Do Next: Stop Burning Tokens, Start Measuring Value
As enterprises expand AI adoption, they are putting serious weight on reducing inference costs and lowering the total cost of ownership across departments. Jacob Bourne notes that “the era of tokenmaxxing is over”—organizations have discovered how easy it is to burn tokens without real value and are pushing back against rising AI bills. That pushback is healthy. It forces providers to offer better price-performance tradeoffs and enterprises to re-evaluate AI spending and ROI based on productivity gains and financial outcomes, not marketing narratives. Price wars will nudge procurement teams to compare proprietary APIs against open-weight alternatives like Kimi K3, which promise greater infrastructure control and lower long-term operating expenses. The conclusion is blunt: in a market where consumer assistants, education tools, healthcare support and productivity apps keep multiplying, the winners inside enterprises will be the teams that treat tokens as a cost center, instrument their workflows, and prove that every AI query earns its keep.






