The uncomfortable truth: aging GPUs are better business than shiny ones
GPU infrastructure economics for AI inference describe how cloud operators earn ongoing returns from graphics processors over many years, balancing power limits, depreciation assumptions, and continuous demand for model serving rather than focusing only on peak training performance or the latest architecture. That definition sounds dry, but the reality is provocative: older hardware is quietly beating new silicon on profitability. CoreWeave’s decision to sign contracts for Nvidia’s A100 GPUs that extend to 2029—nine years after the chips debuted—signals that legacy hardware profitability is not a rounding error on the balance sheet; it is central to AI business models. While the industry obsesses over every new GPU launch, the real money is being made by keeping “good-enough” cards busy in production inference loops, long after the marketing spotlight has moved on.

CoreWeave’s bet: continuous inference turns old GPUs into annuities
CoreWeave’s leadership is openly arguing that the story of AI has shifted from one-off model training to an endless production cycle—training, inference, evaluation, and improvement now form “a single continuous loop.” In this loop, compute is “no longer a one-time requirement concentrated at the beginning of a model's life” but an “ongoing requirement that grows with every application in production and every cycle of improvement.” That framing is crucial for GPU lifecycle management. If inference demand is recurring and rising, the smartest move is not to chase every new architecture, but to keep mature fleets fully utilized for years. Intrator disclosed that CoreWeave has signed A100 contracts running into 2029 and that “pricing for prior generation SKUs is at or above where it was years ago,” a clear sign that legacy AI inference capacity still commands premium revenue.
Power ceilings and legacy halls: why A100s refuse to retire
The most underappreciated driver of legacy hardware profitability is not FLOPS; it is electricity and cooling. An air‑cooled DGX A100 system draws about 6.5kW at maximum load and fits within older data center halls designed for roughly 20kW per rack. Nvidia’s current GB200 and GB300 NVL72 racks, by contrast, pull around 120kW to 140kW and demand direct‑to‑chip liquid cooling—a completely different class of facility. Blackwell‑class systems simply cannot drop into Ampere‑era halls without a rebuild of power delivery and cooling, so the energized, air‑cooled capacity already in place has no higher‑value use. The rational choice is clear: renting nine‑year‑old GPUs is better than leaving those racks dark. This is what GPU infrastructure economics look like in practice: silicon ages, but megawatts and floor space are fixed, and operators squeeze every profitable year out of hardware that matches their existing envelope.
Demand backlog shows mature fleets beat perpetual refresh
CoreWeave’s numbers back up its thesis that continuous AI compute demand favors mature fleets. The company reported quarterly revenue growth of 112 percent year‑on‑year, driven largely by expansion within existing customers rather than new logos. It now carries a substantial revenue backlog and contracted power of 3.7GW in Q2, rising to 4.2GW, while only 1.5GW is currently online—meaning customers have already committed to nearly triple the capacity CoreWeave can deliver today. In this context, GPU lifecycle management becomes about honoring long‑term commitments, not ripping out older cards as soon as a new SKU ships. The CFO noted that capacity coming up for renewal is a limited share of the fleet and that older‑generation average selling prices are at or above levels from a year ago. These are not the economics of rapid obsolescence; they are the economics of infrastructure that pays for itself many times over.
Rethinking the hardware refresh dogma
For years, AI infrastructure narratives treated GPUs as disposable: a Google architect once pegged datacenter GPU service life at one to three years, aligning with a cadence of near‑annual new architectures. Investors worried that faster chip launches would wreck financing models, and one high‑profile critic accused hyperscalers of understating depreciation by extending useful life assumptions. Yet reality is not cooperating with that doom scenario. Nvidia’s CFO has pointed out that A100s sold six years ago are still running at full utilization, while CoreWeave can sign deals that keep 2020 silicon busy through 2029. Inference‑heavy workloads, power‑constrained facilities, and entrenched software stacks all reward stability over speed. The conclusion is uncomfortable but overdue: GPU infrastructure economics now favor keeping older hardware profitably humming in production, and the race to always‑new chips is starting to look more like marketing folklore than sound AI inference strategy.





