MilikMilik

GitHub's AI Traffic Broke Its Infrastructure—and Sent Microsoft to AWS

GitHub's AI Traffic Broke Its Infrastructure—and Sent Microsoft to AWS
Interest|High-Quality Software

What GitHub’s AI Surge Reveals About Modern Infrastructure Limits

GitHub’s recent outages show how AI-driven coding demand can outgrow even large-scale cloud infrastructure plans, forcing platforms to rethink capacity, reliability, and multi-cloud strategies for explosive, unpredictable workloads. The platform has faced months of uneven reliability as AI-assisted coding, agentic workflows, and Copilot usage pushed activity past internal forecasts. GitHub’s May Availability Report admitted to nine incidents that degraded performance, reflecting ongoing GitHub downtime as AI demand reshapes usage patterns. At the same time, GitHub is in the middle of a long-term shift from its own data centers to Microsoft Azure, a move that was supposed to improve cloud infrastructure scaling. Instead, the AI boom exposed Microsoft Azure limitations: capacity expansions that looked generous on paper proved too small in practice, and traditional AI capacity planning models failed to anticipate the speed and intensity of AI-generated code.

From 10x to 30x: When Forecasts Break Under AI Workloads

GitHub’s own numbers show how far its AI capacity planning missed the mark. The company initially targeted a 10x infrastructure increase in October 2025, only to find by February 2026 that a 30x expansion was required to keep up with surging pull requests, commits, and new repositories. According to GitHub SVP Jakub Oleksy, “We’re now serving 40 percent of monolith traffic from Azure (up from 8 percent in February), with Git traffic at 30 percent and repository replication at 99 percent.” Yet outages continued, underlining the gap between planned capacity and AI reality. GitHub reportedly processed about 1 billion commits in an entire year previously; now it sees 1.4 billion commits every month. That kind of step change reflects AI agents and Copilot tools writing, refactoring, and testing code at machine speed, pushing back-end systems into failure modes that earlier models did not account for.

GitHub's AI Traffic Broke Its Infrastructure—and Sent Microsoft to AWS

Why Microsoft Turned to Its Biggest Cloud Rival

The most striking twist in this GitHub infrastructure failure story is Microsoft’s call to Amazon for help. Business Insider reported that Microsoft is adding extra GitHub capacity on Amazon Web Services after AI-driven workloads stressed both GitHub’s legacy systems and its expanding Azure footprint. This is notable because Microsoft has long pitched Azure as the preferred home for AI workloads and had planned to move GitHub fully onto Azure by 2027. Instead, AI coding tools created more commits and repository operations than Azure could absorb in time, exposing practical Microsoft Azure limitations. A Microsoft spokesperson confirmed that GitHub is now using "multiple cloud providers" to secure the elasticity and horizontal scale it needs. In parallel, GitHub even paused new Copilot subscriptions and shifted Copilot to usage-based billing, aligning both infrastructure and economics with the intensity of AI usage.

GitHub's AI Traffic Broke Its Infrastructure—and Sent Microsoft to AWS

Agentic Development: From Autocomplete to Always-On AI Traffic

Traditional autocomplete-style AI, such as early Copilot, generated predictable, short-lived bursts of traffic. The new agentic development wave is different. Coding agents can scan large repositories, plan multi-step changes, modify files, run tests, and open pull requests without human intervention, turning GitHub into a live execution surface rather than a passive code host. Reports note that GitHub has added advanced agents like Claude and Codex into GitHub, GitHub Mobile, and Visual Studio Code for Copilot Pro Plus and Enterprise users, further increasing automated activity. When these agents loop through repositories, they trigger many more Git operations, CI jobs, and metadata updates than human developers working alone. The result is a constant, high-volume stream of AI-generated commits, which created the conditions for GitHub downtime AI demand: database hotspots, network saturation, and cascading failures that GitHub’s engineers are still working to isolate and remove.

Multi-Cloud as a Safety Valve for AI Scaling

GitHub’s move to add AWS capacity on top of Azure highlights a likely pattern for other AI-heavy platforms: cross-cloud partnerships as a safety valve for runaway demand. Microsoft described this as a “multi-cloud strategy” to secure future capacity, elasticity, and horizontal scale. SpaceX and Google’s separate deal for AI compute, alongside Google Cloud’s agreement to sell AI capacity to Anthropic, show that even hyperscalers are becoming each other’s overflow valves. For enterprises, the lesson is clear: AI infrastructure planning that relies on a single cloud and linear growth assumptions is risky. Spikes from new AI features or agent behavior can overwhelm projections overnight. Building for cloud infrastructure scaling now means planning for multi-cloud routing, data portability, and resilience across providers, not only squeezing more headroom out of one stack and hoping AI demand stays within forecasted bounds.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!