What AI token costs are and why they are suddenly a boardroom issue
AI token costs are the metered fees enterprises pay each time an AI model reads or generates text, and as agents and coding assistants run longer, more complex workflows, these token-based charges are exploding and turning AI from a fixed software expense into a volatile, cloud-style consumption bill that finance teams must monitor daily. For years, tools like AI coding assistants were sold on flat subscriptions that roughly matched a human’s workday, but agentic systems have no natural ceiling and can burn through millions of tokens in background loops. Vendors are responding by tying pricing much more tightly to token usage. GitHub Copilot’s shift to GitHub AI Credits, consumed according to underlying API rates, is one prominent example. Enterprise AI pricing is now less about seat count and more about how aggressively teams push models in production workloads.

From flat rate to usage-based: developers discover the cloud bill for AI
The flat-rate era of AI coding tools is ending as vendors move to usage-based billing. GitHub migrated all Copilot plans to GitHub AI Credits, where one credit equals USD 0.01 (approx. RM0.05), and interactions are billed by tokens at each model’s listed API rate. Monthly allowances range from 1,500 credits on Copilot Pro to 20,000 on Max, with business tiers drawing from pooled organisational credits. A preview bill tool and user-level budget controls were added so teams can see projected costs and cap overages. Meanwhile, Anthropic’s Claude Fable 5 arrives inside Copilot at USD 10 (approx. RM46) per million input tokens and USD 50 (approx. RM230) per million output tokens, twice Opus 4.8’s rate. Engineering managers now face the same pattern seen with cloud infrastructure: someone has to watch the dashboard, and token usage budgeting becomes a daily operational task.
Price cuts, losses, and the pressure on OpenAI and Anthropic
As enterprise AI pricing tightens, model providers are squeezed between customer pushback and their own heavy spending. OpenAI is reportedly preparing steep OpenAI price cuts on tokens to win customers back from Anthropic, whose Claude Code has gained momentum with developers. According to financial statements cited by the Financial Times, OpenAI generated USD 13.07 billion (approx. RM60.1 billion) in revenue while incurring a USD 20.92 billion (approx. RM96.2 billion) operating loss, with total costs reaching USD 34 billion (approx. RM156.6 billion). Research and development alone consumed USD 19.18 billion (approx. RM88.3 billion). Those numbers raise hard questions about how sustainable aggressive price cuts are, even as enterprises demand relief from AI token costs. Both companies have already spent billions on infrastructure, and deeper discounts risk pushing margins thinner while usage keeps rising in coding assistants and agents.

Enterprise AI buyers pivot to cheaper models like DeepSeek
With token usage costs climbing, some enterprises are starting to swap premium frontier models for cheaper alternatives. A recent report says Microsoft is considering a self-hosted version of DeepSeek’s V4 model for Copilot Cowork, moving away from costly OpenAI and Anthropic models that are “pricing themselves out of the market.” The motivation is clear: as Copilot Cowork shifts to a metered architecture, every token counts, and cheaper open-source–based models can meaningfully lower AI spend. Other companies have already felt the sting—Uber reportedly blew through its entire AI budget for 2026 in four months after encouraging more internal usage, while some staff engage in “tokenmaxxing” to climb AI-use leaderboards. Not every enterprise is underwater; 8x8, for example, reports that Claude usage for email, feedback analysis, and code is still manageable. But the direction of travel is towards aggressive cost optimisation and model diversity.

How enterprises are rethinking AI token usage budgeting
The combination of soaring AI token costs, shifting enterprise AI pricing, and unpredictable agent workloads is forcing companies to redesign their budgeting and governance. Coding tools that used to be simple line items are turning into cloud-style variable expenses, as seen with Copilot’s usage caps that shut off features when a user’s additional budget is set to zero. This protects the credit card but can halt development work mid-sprint. CIOs are responding by setting tight per-user limits, differentiating between cheap models for quick chat and expensive ones for mission-critical tasks, and monitoring token consumption per team. Some are experimenting with internal policies to discourage tokenmaxxing and encourage prompt discipline. As OpenAI, Anthropic, and rivals enter a price war, enterprises will gain short-term savings, but they will also need sharper visibility into where tokens go and how much value each interaction returns.






