Inference, Not Training, Is Where AI Hits the Wall
The AI inference bottleneck is the point where demand for running large models in real time outstrips the available GPU capacity, forcing companies to ration access, delay new users, or degrade service quality so they can keep existing workloads online and responsive. That is now the industry’s real constraint, and Kimi K3’s launch is the clearest proof so far. Less than 48 hours after release, Moonshot AI had to temporarily pause new subscriptions because usage pushed its GPUs close to their limits, with demand exhausting available capacity even as the company tried to scale up infrastructure. The message is blunt: the hard part of AI is no longer only building frontier models, it is keeping them running when the world decides to use them at scale.

Kimi K3: A 2.8-Trillion-Parameter Stress Test
Kimi K3 went viral precisely because it looks like a frontier model in practice, not just on benchmarks. At 2.8 trillion parameters, it is the largest open-weight model slated for release so far, and it has been trading places with the most advanced systems on independent leaderboards. That capability came with an immediate cost. Within the first 48 hours, usage pushed right up against Moonshot’s ceiling, triggering congestion errors and capacity-related outages before the company drew a hard line and froze new sign-ups to protect paying users. Existing subscribers can continue to use the service while new capacity is added, and future subscriptions will reopen in batches instead of all at once. In other words, performance has overtaken the infrastructure provisioned to support it—and ordinary would-be users are the ones stuck outside the gates.
Agentic Workloads Are Eating GPUs Alive
The capacity crunch around K3 is not a surprise if you look at how people actually use advanced models. As AI shifts from short chats to long coding and agentic workload demands, inference sessions stretch from seconds into minutes or hours, tying up GPU capacity far longer than traditional chatbot queries. Coding flows keep accelerators busy, and agent tasks continuously generate, read, and process tokens during ongoing work rather than ending after a single answer. Citigroup semiconductor analyst Peter Lee warned that as developers build longer agentic workflows, lower per-token inference costs are quickly “re-converted into higher total resource consumption,” with the bottleneck even starting to move from raw compute into server memory. In practice, K3’s cheaper tokens and powerful agent behavior lure more intensive usage, and that usage burns through GPU capacity faster than new infrastructure can be provisioned.
Open Weights, Closed Capacity: The New AI Trade-Off
Kimi K3’s open-weight release is often framed as a win for broader access, but the subscription freeze exposes the uncomfortable catch. Open weights do not mean open capacity: whoever hosts the model still has to pay the inference bill and live within GPU capacity constraints. The token economics make this palpable. Moonshot charges $3 per million input tokens and $15 per million output tokens for K3, roughly 40% cheaper than one rival model and around 70% cheaper than another. That pricing invites heavy, agentic workload demands, yet the infrastructure lags. Meanwhile, the scramble for computing power has triggered a data-center construction boom among major cloud providers, with tens of billions committed to AI and cloud infrastructure over the coming years. The value curve is bending away from the glamorous model layer and toward chipmakers, cloud platforms, and the teams who handle AI infrastructure scaling.
Designing for Scarcity: How AI Builders Must Respond
Moonshot’s response—pausing new sign-ups, prioritizing existing subscribers, and planning a split between Kimi Membership for general use and a separate Kimi Code Membership with dedicated capacity—shows where AI operators are headed. They are designing product tiers and rationing access because infinite, cheap API usage is no longer realistic. For ordinary users, this means you may be turned away even when you are ready to pay, as providers choose reliability for current workloads over unfettered growth. For companies, it is a wake-up call: model performance can no longer be considered in isolation from infrastructure costs and availability. Longer agentic workflows and lower token prices are attractive, but they drive total resource consumption sharply higher. The new competitive advantage is not only having a capable model—it is architecting systems that stay online when demand spikes, even if that means saying “no” to new customers today so the ones you already have do not watch your service collapse.






