From Tokenmaxxing to Thrift-Maxxing: What AI Budget Rationalization Really Means
Enterprise AI budget rationalization is the shift from routing most workloads to a single premium model for speculative innovation toward selectively using multiple cheaper, fit-for-purpose models based on measurable return on investment and cost-per-inference. It replaces hype-driven spending with deliberate choices about which models, platforms, and infrastructure are needed for specific business tasks, and which expensive capabilities can be cut without hurting outcomes. That shift is now showing up directly in enterprise AI spending cuts: analysis of 2.4 billion API calls shows token costs down 67 percent year-over-year, with blended cost per million tokens falling from USD 18.40 (approx. RM84.40) to USD 6.07 (approx. RM27.80). The headline story is simple: enterprises have decided they do not have to blow their budgets on AI anymore. After a year of “tokenmaxxing”, buying the biggest models for every task, companies are pivoting to “thrift-maxxing”, mixing cheaper providers and open-source models with familiar names. This is not retreat; it is a correction. AI is moving from a prestige project to a disciplined line item where enterprise AI ROI must be proven, not assumed.
Multi-Model AI Platforms Are the New Default, and Single-Vendor Lock-In Is the First Casualty
The first thing to go under scrutiny is single-vendor lock-in. Enterprises are abandoning one-provider AI strategies and shifting to multi-model AI platforms as new data shows cost savings above 60 percent and deployment timelines cut by two-thirds. Average models per enterprise account have jumped from 2.1 to 4.7, making multi-model architecture the default rather than an experiment. On unified aggregation layers that connect providers such as OpenAI, Google, Anthropic, xAI, DeepSeek, Alibaba, ByteDance, and MiniMax, intelligent routing distributes work across premium, mid-tier, and low-cost models. In Q1 2025, 73 percent of enterprise token volume went to the two most expensive model tiers; by Q1 2026 that share sank to 31 percent, with 69 percent flowing to mid-tier and cost-efficient models matched to task complexity. The most expensive frontier models are no longer the default option for mundane tasks, because they do not earn their price. According to AI.cc, “multi-model strategy is no longer optional… businesses still routing all AI requests to a single premium provider are overpaying by a significant margin”. That is the clearest sign yet that multi-model AI platforms are not a side show; they are now the core procurement strategy.
Open-Source and Cheaper Models Are Eating the Cost Curve
If single-vendor lock-in is the first cut, overpaying for frontier models is the second. Fed up with ballooning costs, companies large and small are starting to use lower-priced models, including ones built in China, and add them alongside high-profile products from OpenAI and Anthropic, shopping a la carte for their AI needs. The logic is blunt: the most powerful and expensive models are not necessary for relatively mundane tasks like simple classification or structured data extraction. Data from AI.cc shows how hard this is hitting the cost curve. Enterprise teams were previously sending simple work through frontier models just because those were already integrated. Now, intelligent routing alone accounts for 34 percentage points of the total 67 percent cost reduction. Open-source models have accelerated that change: they went from 11 percent of enterprise token volume in Q1 2025 to 38 percent in Q1 2026, a 245 percent share increase driven by aggressive pricing from providers such as DeepSeek. Enterprises using multi-model routing report median token cost reductions of 71 percent, with the top quartile exceeding 80 percent. This is a brutal message for premium labs: the margin they enjoyed on “tokenmaxxing” is now directly exposed to cheaper competitors and open-source alternatives.
What Gets Cut Inside the Stack: Infrastructure, Overkill Models, and Slow Integrations
The spending cuts are not vague belt-tightening; they are targeted at three obvious waste zones. First, overbuilt infrastructure built around one frontier model is being reconsidered, because routing every workload to a single premium provider is no longer cost-effective. Token costs have already been pushed down from USD 18.40 (approx. RM84.40) per million tokens to USD 6.07 (approx. RM27.80) through smarter routing. Second, unnecessary use of frontier models is being stripped away. Work that can be handled by mid-tier or open-source models is moved there, cutting the share of volume hitting the priciest tiers from 73 percent to 31 percent year-on-year. Third, slow custom integrations are being replaced with faster multi-model layers. Teams using multi-model infrastructure deploy production AI agents in a median of 3.6 weeks, versus 11.2 weeks for single-provider integrations—a threefold improvement in time to market. That speed gain is not a nice-to-have; it makes AI projects more likely to hit deadlines and translate into enterprise AI ROI instead of sitting in pilot limbo. In short, companies are cutting sunk-cost vanity projects and keeping anything that proves it can deliver cheaper inferences, faster deployment, or clearer business outcomes.
The Maturing Enterprise AI Playbook: Pragmatic Deployment over Hype
The shift to enterprise AI spending cuts is not a sign that AI is fading; it is evidence that AI is finally being treated like any other technology investment. As multi-model routing, prompt caching, and aggregated pricing redefine the economics of AI APIs, organizations that fail to adapt risk falling behind competitors that have already secured these efficiencies. US companies have flipped from “tokenmaxxing” to “thrift-maxxing”, mixing cheaper Chinese models with OpenAI’s and Anthropic’s products and in the process threatening the valuations of the frontier labs. This is what a maturing market looks like. Hype-driven procurement gives way to multi-model AI platforms, careful cost-per-inference tracking, and a bias for open-source where it performs well. Premium models will still have a role for complex, high-stakes tasks, but they must earn that role instead of being default choices. The winners in this new era will not be the companies that shouted the loudest about AI, but those that rebuilt their stacks to use the right model for the right job at the right price—and cut everything else.






