AI data centers: where the field actually stands
Frontier clusters went from 150 MW in 2023 to gigawatt-class in 2025, and the binding constraint is now power rather than chips. Four hyperscalers guide over $300B a year combined — the fastest capex ramp in corporate history.
AI data centers carry four metrics and 7 milestones here. This is the youngest field on the tracker and the one where the numbers move fastest.
Power is the headline metric, because power rather than chips is now the binding constraint on frontier training. xAI's Colossus in Memphis and Meta's Prometheus in Ohio each reach about 1,000 MW, with Microsoft's Fairwater and OpenAI sites near 900 MW, Google's TPU campuses near 800 and Oracle's Stargate Abilene phase one near 700. The 2023 baseline for a frontier cluster was 150 MW. That is a roughly sevenfold jump in two years, and announced future campuses reach toward 10 GW.
GPU count is the coarser but more vivid measure. Colossus wires about 200,000 accelerators, Microsoft's OpenAI training sites about 150,000, Stargate Abilene about 130,000 and Meta's largest Llama cluster about 100,000 — against a 2023 frontier baseline of 25,000. Counts mix GPU generations, H100 alongside GB200, which is why the tracker reads them alongside power rather than on their own.
Capital expenditure shows the scale of the bet. Amazon guides about $100B a year, Microsoft $80B, Google $75B and Meta $65B — over $300B combined, the fastest capex ramp in corporate history. Combined hyperscaler capex crossed $200B a year back in December 2024.
The milestone record is short because the field is young. xAI brought the first six-figure GPU cluster online in September 2024, roughly 100,000 H100s at about 150 MW, built in months. The Stargate Project was announced in January 2025: $500B and roughly 10 GW, with the first multi-gigawatt campus rising in Abilene, Texas — the largest compute buildout ever committed. Colossus then scaled to 200,000 GPUs and gigawatt-class power as GB200 racks came online, the first single site to approach 1 GW of AI compute.
The most structurally interesting 2026-adjacent result is AWS's Project Rainier, activated in October 2025: nearly 500,000 of Amazon's own Trainium2 chips across multiple US data centers, one of the world's largest AI clusters and the biggest built on custom silicon rather than Nvidia GPUs. Anthropic trains and serves Claude on it, at more than five times its previous training compute, scaling toward a million-plus Trainium2 chips. That is a proof point that frontier-scale compute can run on a hyperscaler's in-house accelerators — a meaningful check on single-vendor dependence.
Two things remain ahead. The race to a single million-GPU cluster and multi-gigawatt campuses is under construction, not operational. And the locked milestone — a single campus delivering on the order of 10 GW — is targeted around 2029. Research output on data-center and AI-infrastructure efficiency runs about 2,227 papers a year, modest against the scale of spending, which is itself worth noting: this is a field being built far faster than it is being studied.
- Latest
- 1 GW MW
- 2025
- Goal
- 10 GW MW
- Stargate full build-out (~2029)
Who's involved
Elon Musk's AI company; builds the Grok models and the large Colossus GPU cluster.
Maker of the open-weight Llama models; Llama 3.1 405B was trained at 3.8e25 FLOP.
Azure Quantum's topological-qubit approach; unveiled Majorana 1, a processor built on Majorana zero modes, in 2025 with a million-qubit roadmap.
Data-center partner in the Stargate project; building the first multi-GW campus in Abilene, Texas.
The world's largest cloud provider; the biggest AI-infrastructure spender, guiding ~$100B capex in 2025.
Reader response
Reactions and comments are reader opinion — unverified, and never part of the newsroom's fact record.
Comments 0
No comments yet — be the first.