AI infrastructure, from first principles
Behind every chatbot answer is a chain of physical infrastructure. This track follows it from the chip in the rack to the transmission system feeding the campus, using tools you can operate rather than explanations you only read.
Size an AI facility in your head, from a 1 kW chip to a grid-scale campus · read PUE and capacity claims critically · explain why the grid, not the silicon, sets the timeline.
What actually happens when you ask an AI a question
Your prompt travels to a data center, where it reaches a cluster of accelerators: chips designed for the enormous volume of arithmetic AI models require. Generating a response takes computation, and computation takes electricity. The answer on the screen therefore has a physical path through generation, transmission, transformation, distribution, and silicon.
Energy-per-query estimates vary with the model, hardware, response length, batching, and facility efficiency. A single text response may require only a fraction of a watt-hour, but small units become material when multiplied across large volumes, and inference sits on top of the much larger, concentrated demand of training.
Two activities dominate. Training teaches a model by occupying large accelerator clusters for weeks or months. Inference runs the finished model for users, continuously and across many locations. The two workloads can stress infrastructure differently.
Keep one mental model: AI infrastructure is a chain of containers, each often about an order of magnitude larger than the last. Chip, server, rack, pod, hall, building, campus, grid. Once the chain is visible, headlines about gigawatts and interconnection queues become physical rather than abstract.
From one chip to the grid, in eight steps
Click a rung. The blue dot shows where you are on a logarithmic power scale: each step is roughly a factor of ten.
The words the industry actually uses
Filter by level, tap a term. Basics assume nothing; Advanced assumes you've read the ladder.
Two tools for how this behaves, not just what it is
Cooling is the technology transition inside the transition. Every watt a chip consumes becomes heat that must leave the building. The industry is climbing a cooling ladder of its own. Tap through it:
Efficiency has one headline number. PUE: total power over IT power. Drag the slider and watch where a fixed megawatt of compute takes the whole facility.
Where a megawatt goes
PUE 1.40Overhead split (≈75% cooling, 15% distribution losses, 10% other) is an illustrative industry-typical breakdown; real splits vary by climate and design. Fleet and industry PUE reference points from operator disclosures and Uptime Institute surveys.
Then behavior. Utilities plan around load shapes, and AI brought two new ones at once. Toggle below: training clusters move in synchronized steps (thousands of GPUs accelerating and pausing together, swinging load in seconds) while inference follows the gentler daily rhythm of human activity. Grids historically saw nothing like the first curve from a customer.
Two loads, two personalities
Training: near-full utilization punctuated by synchronized dips: checkpoints, data loading, failures and restarts. Swings of this kind (documented by grid reliability bodies studying large synchronized loads) are why ramp-rate language now appears in large-load interconnection agreements.
One more asymmetry worth knowing. Training can tolerate interruption because jobs restart from checkpoints, so some operators accept leaner redundancy in exchange for speed and cost. Inference is customer-facing and lives behind stricter uptime promises, with the fuller N+1 or 2N electrical topologies to match. Same chips, different obligations, different buildings.
Common misconceptions
Developers announce far more than gets built: the same project may be shopped to several utilities, and interconnection queues are full of placeholders. Energized megawatts, not press releases, are the number to track.
PUE only measures facility overhead. A building full of idle GPUs at PUE 1.1 wastes vastly more energy than a busy one at 1.4. Useful work per watt is a different (and harder) question.
It depends entirely on cooling design. Evaporative systems do consume significant water; closed-loop liquid and air-cooled designs consume very little on site, sometimes trading water use for higher power use instead. There is no single answer.
Energy is rarely the binding limit: delivery is. The constraint is almost always transmission capacity, transformer lead times, and interconnection process, not generation fuel. That's why the queue, not the power plant, sets the schedule.
Four questions before you go
Where to go next
Sources & methodology
Figures on this page are order-of-magnitude teaching values, hedged where designs vary. Primary reference points: vendor system specifications (accelerator TDPs, 8-GPU system power, 72-GPU rack-scale systems); Uptime Institute annual Global Data Center Survey (industry PUE, rack densities); operator sustainability disclosures (fleet PUE); The Green Grid / ISO-IEC 30134-2 (PUE definition); EIA residential consumption data (megawatt-to-homes intuition); IEA Energy & AI (demand outlook); DOE and Lawrence Berkeley National Laboratory studies (transmission and interconnection timelines); grid reliability bodies’ work on large synchronized loads (load-shape behavior); independent energy-per-query estimates (fraction-of-a-watt-hour figure). Load-profile curves are illustrative shapes, not measured traces.
ENVIZN