Data Center Infrastructure: A Practical Reference

A decade of data center design assumed 5–10 kW racks and air. AI training racks draw 80–140 kW and are heading past 250, which breaks the air assumption, the power assumption, and usually the building. This guide catalogs 35 systems across seven classes, with the rack density each one actually supports, how much water it uses, and whether it can go into a hall that already exists.

35systems
7classes
8families
DensityRack power density the system can support. Low under 10 kW/rack is legacy enterprise · Medium 10–40 kW covers most cloud racks · High 40–100 kW is where air stops working · Extreme above 100 kW is AI training. Most systems span a band, and the top of that band is what matters.Each entry covers a span of bands, and picking several widens the results.
Used inWhere the system is normally deployed. AI training is called out separately from hyperscale because its density, power ramp rate, and cooling requirements differ enough to change the design.Pick several tags and an entry has to carry all of them, so each one narrows the results.
Water useOn-site water consumption when the system is running. This is the metric communities and regulators ask about, and it trades directly against electricity: evaporative cooling saves power and spends water, closed loops do the reverse. Untagged where the system does not touch the water balance.Each entry sits in exactly one band, so picking several widens the results.
RetrofitHow disruptive the system is to install in a facility that already exists. Drop-in = goes into a live hall rack by rack · Hall retrofit = needs a room taken out of service and reworked · New build = the building has to be designed around it.Each entry sits in exactly one band, so picking several widens the results.
MaturityMature = standard practice with many suppliers · Established = commercial and widely deployed but newer · Emerging = shipping at scale only in the last few years · Early = deployed in pilots and single sites.Each entry sits in exactly one band, so picking several widens the results.
Class I

Power delivery

from the utility takeoff to the rack outlet4 systems

Everything upstream of the building starts here: a transmission or subtransmission takeoff, an on-site substation with transformers stepping 69–230 kV down to medium voltage, protection, metering, and usually a ring or radial medium-voltage distribution around the campus. A modern hyperscale campus takes 100–1,000 MW, which is a small city's worth of load arriving at one meter, and that is why site selection now begins with a conversation about the substation rather than about fiber or land.

Strengths & weaknesses

Utility service is by far the cheapest power available, at $0.04–0.10/kWh in most US markets, and it comes with no on-site fuel logistics or emissions permit. Reliability from a well-designed dual feed is high. The weaknesses are all schedule and scale. Interconnection studies for a large load run one to four years, the substation transformers behind them run 80–144 weeks, and in constrained regions the answer is simply no until a transmission upgrade lands. Utilities also increasingly ask for curtailment commitments, contribution in aid of construction, or both, in exchange for a faster connection.

When to use

Utility service is the default and should be the plan for any facility that will run for a decade. The decisions are where and on what terms. Choose sites with existing capacity headroom rather than sites that look good on land price, because the interconnection queue dominates the schedule. Where the queue is long, expect to negotiate flexible-load terms, which can move a project years earlier. On-site generation is a bridge or a supplement to this, not a replacement, unless the facility is genuinely remote.

Key numbers

Hyperscale campuses take 100–1,000 MW · service typically at 69–230 kV stepping to 13.8–34.5 kV on site · US industrial power commonly $0.04–0.10/kWh · large-load interconnection studies take one to four years · substation transformer lead times of 80–144 weeks · US data centers used about 176 TWh in 2023, roughly 4.4% of national electricity.

Examples

The northern Virginia cluster, where Dominion's transmission constraints have become the limiting factor on new capacity; ERCOT's large-load interconnection process, built specifically around data centers and crypto; the Ohio and Georgia campuses sited primarily on available substation capacity.

Economic profile

Power cost is 15–30% of a data center's total cost of ownership, so a two-cent difference in rate matters, but availability matters more. Developers now pay premiums for "powered land," a site with an executed interconnection agreement, because that agreement is worth more than the acreage. Where the utility asks for contribution in aid of construction toward network upgrades, that becomes a real capital line item and is worth negotiating alongside the rate.

Videos
Data Center Power Flow: From Utility Grid to Server RackMEP Academy · 50k+ views
Data Center Power Chain - AnimationTechTrainerNJ · 100k+ views
Data Center Power Explained (It's simpler than you think)Base Config · 10k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · Key Questions on Energy and AI (IEA)

Below the UPS, power reaches the racks one of two ways. The traditional route runs cables from a power distribution unit through conduit or under a raised floor to each rack, one circuit at a time. The alternative is overhead busway: a continuous enclosed bar running above the rows, into which tap-off boxes plug anywhere along its length. Busway turns adding a circuit from an electrician's project into a plug-in operation, which is why nearly every hall built for changing IT load now uses it.

Strengths & weaknesses

Busway's advantages are flexibility and airflow. Circuits move as racks move, capacity is added without pulling cable, and getting power out of the underfloor plenum leaves the plenum for air. Metering at the tap-off gives per-rack visibility for free. Against that, busway costs more up front than conduit for a fixed layout, its rating has to be chosen at design time and is expensive to change, and a busway sized for 5 kW racks is now the constraint in many older halls. Cable is cheaper and entirely adequate when the layout will not change.

When to use

Use overhead busway in any hall where rack configurations change, in colocation where tenants come and go, and wherever density is rising, since it is the cheapest part of the chain to over-size. Size it for the density you expect at end of life rather than at day one, because replacing busway means taking the row down. Stay with fixed cable distribution for a small enterprise room with a stable layout, and for retrofits where working overhead is not practical. Whichever route, put revenue-grade metering as far downstream as budget allows; per-rack data is what makes capacity planning possible.

Key numbers

Busway ratings commonly 250–1,200 A, chosen at design and hard to change · tap-off boxes add a circuit in minutes without an outage · per-tap metering is standard on modern systems · overhead routing frees the underfloor plenum for air · cable distribution is cheaper for a fixed layout and worse for a changing one.

Examples

Overhead busway in nearly all colocation halls built since about 2010; underfloor cable distribution in legacy enterprise rooms and in halls that predate high density; busway rating limits now driving hall rebuilds as racks move from 5 kW to 30 kW.

Videos
Is a Busway System Right for Your Data Center?Anixter · 1k+ views
Lesson 7 - Part 2: Power Distribution for Data Centers and UPSEngineering and Donuts · 10k+ views
Further reading

Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP)

A rack power distribution unit is the strip inside the cabinet that turns one or two feeds into the dozens of outlets the servers plug into. Four grades exist: basic, which is a strip; metered, which reports total draw; monitored per-outlet, which reports each outlet; and switched, which can turn outlets on and off remotely. Racks are normally fed from two independent PDUs on separate upstream paths, so either can be lost without dropping dual-corded equipment.

Strengths & weaknesses

The metering is the value. Per-outlet data tells you what is actually drawing power, which is the input to every capacity and stranded-capacity decision, and switched outlets allow a remote power cycle instead of a site visit. They are cheap, install in minutes, and need no design change. The weaknesses are that they are one more thing in a hot cabinet with a finite life, that a failed PDU takes out one power path, and that their built-in circuit breakers trip on inrush if a whole rack is energized at once. At very high density they run out of runway: a 130 kW rack cannot be fed by conventional 208 V strips.

When to use

Use metered or per-outlet PDUs everywhere, since the incremental cost over basic is small and the visibility is what capacity planning runs on. Use switched units in colocation, at the edge, and anywhere staff are not on site. Above about 40–50 kW per rack, stop and look at higher-voltage distribution and rack power shelves instead, because the number of cords and the copper needed to feed the rack conventionally becomes impractical. Always confirm the PDU's own breaker curve against the servers' inrush before a full-rack power-up.

Key numbers

Typical rack feeds 208 V or 415 V three-phase at 30–60 A, giving roughly 5–20 kW per PDU · dual-fed racks are standard, so each path carries the full load in a failure · per-outlet metering typically within plus or minus 1% · switched units allow remote power cycling · conventional strips run out somewhere around 40–50 kW per rack.

Examples

Metered and switched PDUs from Vertiv, Raritan, APC, and ServerTech in nearly every colocation cabinet; per-outlet data feeding DCIM capacity models; 415 V distribution in hyperscale halls, which delivers 240 V line-to-neutral to equipment and cuts conductor size.

Videos
Understanding Rack Power Distribution (PDU)Critical Facilities Connect · under 1k views
Rack Power (PDU) terms and technologyTechTrainerNJ · 100k+ views
Power Distribution Units | Data Center Rack Power Distribution Units | VueNowVueNow Official · 5k+ views
Further reading

Data Center Energy Efficiency Toolkit (Berkeley Lab Center of Expertise)

Conventional data center power converts several times: AC in, DC in the UPS, AC out, then AC to DC again in every server supply. High-voltage DC distribution cuts the middle out. Power is rectified once at the room or row level and distributed as DC, historically at 380–400 V and now increasingly at plus or minus 400 V DC for AI racks, down a busbar into the rack where power shelves feed the servers. Batteries connect directly to the DC bus, so there is no inverter between them and the load.

Strengths & weaknesses

Removing conversion stages saves a few percent of total facility power and removes hardware that can fail. At AI rack densities the bigger argument is copper: at 130 kW a rack fed at 415 V AC needs an impractical number of large conductors, and raising the voltage is the only way to keep the busbar a reasonable size. Batteries on the DC bus also ride through faster than an inverter can. The weaknesses are ecosystem and safety. DC at 400 V does not self-extinguish an arc the way AC does, so connectors and protection are specialized, the supply chain is thin compared with AC, and equipment has to be bought for it rather than adapted.

When to use

Consider it for new AI training halls at 100 kW or more per rack, where the copper argument alone justifies it and the whole hall can be designed around one architecture. It also fits telecom-adjacent facilities already comfortable with 48 V DC practice. Do not retrofit it into a mixed hall, since running two distribution architectures doubles the spares and the training. For anything under about 50 kW per rack, conventional AC distribution is cheaper, better supported, and efficient enough that the conversion savings do not repay the disruption.

Key numbers

Distribution at 380–400 V DC, and plus or minus 400 V DC in recent AI rack designs · removes two conversion stages, saving roughly 2–5% of facility power · batteries connect directly to the bus with no inverter · DC arcs do not self-extinguish, so protection and connectors are specialized · adopted mainly in new-build AI halls rather than retrofits.

Examples

Open Compute Project rack power designs, which moved from 12 V to 48 V bus bars and then toward higher DC distribution; NVIDIA's 800 V DC reference architecture for high-density AI racks; long-standing 48 V DC practice in telecom central offices, which is the same idea at lower voltage.

Videos
Power Distribution for AI Data Centers | Schneider ElectricSchneider Electric · 1k+ views
Further reading

Electrical Efficiency Measurement for Data Centers, White Paper 154 (Schneider Electric)

Class II

Standby generation

engines that start when the grid fails2 systems

Standby diesel generators are what carry a data center through a utility outage longer than the batteries can. A 2–3 MW unit in an outdoor enclosure or generator room starts on a signal from the transfer switch, reaches rated speed and voltage in about 10 seconds, and picks up load through an automatic transfer switch. A large campus has dozens of them in N+1 or 2N arrangements, plus fuel storage sized for 24–72 hours of full-load running and a contract for resupply.

Strengths & weaknesses

Diesel generation is the most proven backup there is: fuel is dense and storable, starting is reliable when the maintenance is done, and the units run for years at very low duty. Nothing else combines that availability with on-site fuel. The costs are permitting and testing. Generators are permitted air emission sources with hour limits, and monthly load-bank testing burns fuel and produces the emissions everybody notices. Fuel goes stale and needs polishing, wet stacking damages engines run lightly, and in constrained air-quality districts a large generator plant can be the hardest permit on the project.

When to use

Diesel remains the default for standby duty and should be assumed unless there is a specific reason not to. Sizing follows the redundancy target: N+1 for concurrent maintainability, 2N where the design has to survive a failure during maintenance. Look at alternatives where air permits are the binding constraint, where the site cannot store fuel, or where the utility will pay for the capacity: in some markets, generators enrolled in a demand-response program earn enough to change the business case. Batteries and fuel cells substitute for part of the duty but not for a multi-day outage.

Key numbers

Typical unit 2–3 MW, with dozens per campus · start to full load in about 10 seconds · fuel stored for 24–72 hours at full load · monthly testing under load is standard practice · permitted run hours are usually the binding regulatory constraint · capital roughly $500–900 per kW installed.

Examples

Generator yards at every large colocation and hyperscale campus; Manhattan data centers with generators on the roof and fuel in the basement; Irish and Dutch facilities where generator air permits became a public planning fight; demand-response programs that pay generator fleets to run at peak.

Videos
A DAY in the LIFE of the DATA CENTRE | GENERATOR TESTING with ASH!Custodian Data Centres · 100k+ views
What is a Generator and How It Works in a Data Center 1080pCoreSite · 10k+ views
Cat® Diesel Generator Sets Supply Emergency Power to Manhattan’s Data CenterCat Electric Power · 100k+ views
Further reading

Uptime Institute Global Data Center Survey 2025 (Uptime Institute)

Prime power means the site generates most of its own electricity continuously rather than waiting for an outage. For data centers that currently means reciprocating gas engines or aeroderivative gas turbines behind the meter, sized in the tens to hundreds of megawatts, often with a grid connection retained as backup rather than as the primary supply. The driver is not cost; it is that the interconnection queue in the target market is longer than the business can wait, and gas generation can be permitted and built in 12–24 months.

Strengths & weaknesses

Speed is the product: gas engines arrive as containerized modules and a plant can be running long before a transmission upgrade would finish. The site controls its own capacity and is not exposed to a utility's schedule. Against that, delivered electricity costs more than utility service in most markets once fuel, maintenance, and capital are counted, the site now runs a power plant with the staffing that implies, gas supply needs a pipeline connection with its own lead time, and the air permit for continuous operation is far harder than for standby. Emissions are a live public issue in every jurisdiction where this has been tried.

When to use

Consider prime power where the alternative is waiting years for interconnection and the compute has to be online sooner, and where gas supply and air permits are genuinely obtainable. It also works as a bridge: run on gas now, connect to the grid when the upgrade lands, and keep the plant for backup and peak shaving. Do not choose it where power price drives the model, where the operator has no appetite to run generation, or where the site's emissions profile is a reputational problem. Fuel cells cover a similar niche with lower emissions and higher cost.

Key numbers

Gas engine plants of 10–500 MW, built in 12–24 months against multi-year interconnection waits · reciprocating engines around 40–45% efficient, aeroderivative turbines similar in simple cycle · delivered cost typically above utility rates once fuel and capital are counted · continuous-run air permits are far more restrictive than standby permits · gas pipeline connection carries its own multi-year lead time.

Examples

Texas data centers built with behind-the-meter gas generation while awaiting ERCOT interconnection; xAI's Memphis site, whose gas turbines drew air-permit scrutiny; VoltaGrid and similar modular gas fleets marketed specifically for bridge power; several announced projects pairing gas today with a grid connection later.

Videos
Inside Climate News: Data Centers Are Building Their Own Gas Power Plants in TexasKXAN · 5k+ views
Powering a Data Center Off the Grid: Part 1 – Natural Gas SolutionsVoltaGrid · under 1k views
VoltaGrid's Data Center SolutionVoltaGrid · 5k+ views
Further reading

Key Questions on Energy and AI (IEA) · AI: Five charts that put data-centre energy use - and emissions - into context (Carbon Brief)

Class II

UPS & ride-through

carrying the load through the gap4 systems

A double-conversion uninterruptible power supply rectifies incoming AC to DC, holds a battery on that DC bus, and inverts back to AC for the load. Because the load is always fed from the inverter, the utility waveform never reaches the IT equipment: sags, harmonics, and frequency wander are all filtered out, and when the utility fails the batteries simply keep supplying the same DC bus with no transfer at all. Ride-through is typically 5–15 minutes, which is far longer than the 10 seconds a generator needs.

Strengths & weaknesses

It is the cleanest power available and the transfer is genuinely seamless. Modern units run at 96–97% efficiency in double conversion, and above 98% in eco mode where the load runs on utility power with the inverter standing by. The costs are capital, footprint, and the batteries. A UPS plant plus batteries is a large room, valve-regulated lead-acid strings need replacing every 3–5 years, and eco mode trades a little transfer risk for the last point of efficiency. Every conversion stage is also a stage that can fail, which is why UPS modules are deployed N+1.

When to use

Double conversion is the default for anything that cannot tolerate a momentary interruption, which is most IT load. Choose it wherever utility power quality is poor, since the filtering is worth as much as the ride-through. Consider eco or multi-mode operation to recover efficiency on a clean supply, with the transfer time verified against the equipment's own tolerance. Where the load is small and the power is clean, a line-interactive unit is cheaper. Where the site wants the batteries to earn money between outages, a lithium battery plant with grid-services capability is the better structure.

Key numbers

Efficiency 96–97% in double conversion, above 98% in eco mode · ride-through typically 5–15 minutes at full load · no transfer time, because the load never leaves the inverter · valve-regulated lead-acid strings last 3–5 years, lithium 8–10 · deployed N+1 at module level in most designs.

Examples

Modular UPS systems from Vertiv, Schneider, Eaton, and ABB in almost every colocation facility; eco-mode operation now common in hyperscale, where power quality is good and the efficiency point is worth chasing; Uptime Institute survey data showing average PUE stuck near 1.5, of which the UPS is a small but persistent contributor.

Videos
How Data Center UPS Systems WorkMEP Academy · 10k+ views
Uninterrupted Power Supply (UPS) Operating modesRockz Automation · 50k+ views
What is a UPS? (Uninterruptible Power Supply)RealPars · 500k+ views
Further reading

Electrical Efficiency Measurement for Data Centers, White Paper 154 (Schneider Electric)

A flywheel UPS stores energy as rotation instead of chemistry. A steel or composite rotor spins in a low-friction bearing, usually magnetically levitated in a partial vacuum, and on a power failure the motor becomes a generator and delivers 15–30 seconds of full-load power. That is enough to start a generator, which is all the ride-through most designs need. A diesel rotary UPS goes further and puts the flywheel, a motor-generator, and a diesel engine on one shaft, so the same machine conditions power, rides through, and then runs on fuel.

Strengths & weaknesses

No batteries is the point: nothing to replace on a five-year cycle, no thermal runaway risk, no dedicated battery room with its own cooling and fire suppression, and a 20-year life. Footprint is small for the power and the units tolerate heat that would shorten battery life. Against that, ride-through is seconds rather than minutes, so the design depends completely on the generator starting; the rotating machine needs mechanical maintenance a static UPS does not; and standby losses of 1–2% run continuously. Fewer vendors serve the market than for static systems.

When to use

Choose flywheel or rotary where generators are reliable and tested, where battery replacement cost and room space are real burdens, and where the site is hot enough that batteries would age quickly. Diesel rotary suits large single-block loads and has a strong record in Europe. Do not choose it where the design must survive a generator start failure, since seconds of ride-through leaves no second chance, and do not choose it if the operator wants the stored energy to do anything besides ride-through. Lithium batteries can also provide grid services; a flywheel cannot.

Key numbers

Ride-through 15–30 seconds at full load, against 5–15 minutes for batteries · rotor life around 20 years with no scheduled replacement · standby loss roughly 1–2% of rating · no battery room, no thermal runaway risk · tolerates ambient temperatures that would shorten battery life significantly.

Examples

Active Power and Vycon flywheel systems in North American facilities; Hitec and Piller diesel rotary UPS installations across European data centers; Cisco's Texas facility, an early large rotary deployment; flywheels used as a bridge alongside batteries in hybrid designs.

Videos
Data Center World: Flywheel UPS DemonstrationData Center Knowledge · 50k+ views
The Cat Flywheel UPSPeterson Cat · 5k+ views
Dynamic Rotary UPS at Cisco's Allen TX Data Centercycloneinteractive · 50k+ views
Further reading

Uptime Institute Reports (Uptime Institute)

Lithium-ion has largely displaced valve-regulated lead-acid as the energy store behind the UPS. The cells sit in rack-mounted cabinets with a battery management system that monitors every module, and the chemistry is usually lithium iron phosphate rather than the nickel-rich chemistries used in vehicles, because thermal stability matters more than energy density when the pack lives in a building full of servers. Once the plant is large enough, it stops being only a UPS: the same batteries can shave peaks, respond to grid frequency, or shift load.

Strengths & weaknesses

Compared with lead-acid it takes about a third of the footprint and a quarter of the weight for the same energy, lasts 8–10 years instead of 3–5, tolerates higher ambient temperature, and reports its own state of health. Over a 10-year life the total cost is usually lower despite a higher purchase price. The weaknesses are code and fire. Lithium installations face specific requirements under NFPA 855 and local fire codes, including spacing, detection, and sometimes deflagration venting, and the permitting conversation is materially harder than for lead-acid. Cell supply is also exposed to the same market as vehicles.

When to use

Use lithium for any new UPS energy store where footprint or replacement labor matters, which is nearly all of them, and specify lithium iron phosphate unless there is a specific reason for a nickel chemistry. Size beyond ride-through if the local market pays for demand response or frequency service, because the incremental cells are cheap relative to the rest of the plant. Engage the fire marshal early rather than late. Keep lead-acid where an existing room and its suppression are sized for it and the replacement cycle is already funded.

Key numbers

Roughly a third the footprint and a quarter the weight of lead-acid for the same energy · service life 8–10 years against 3–5 · tolerates higher ambient temperature, so the battery room can run warmer · governed by NFPA 855 and local fire code, which drives spacing and detection · large plants can also bid into demand response and frequency markets.

Examples

Lithium iron phosphate UPS cabinets now standard in new hyperscale builds; Microsoft and Google installations using UPS batteries for grid services; Irish and Dutch facilities offering battery capacity to system operators in exchange for faster connection.

Videos
What's New with UPS Batteries?Eaton · 1k+ views
ORR Protection Lithium-Ion Battery Q&A: Data Center Code ComplianceORR Protection · under 1k views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab)

A fuel cell converts fuel to electricity electrochemically rather than by combustion. In data centers the dominant type is the solid oxide fuel cell running on natural gas or biogas, delivered as containerized modules of a few hundred kilowatts and stacked to whatever capacity the site needs. Because they run continuously rather than on standby, fuel cells are prime power with a different emissions profile, not a generator substitute. Proton-exchange membrane cells running on hydrogen have been demonstrated in the backup role but remain rare.

Strengths & weaknesses

Electrical efficiency of 50–60% beats a reciprocating engine, and because there is no combustion the NOx and particulate emissions are very low, which is what makes them permittable where engines are not. Modules install quickly and scale in small increments. The costs are capital and fuel. Installed cost per kW runs several times a gas engine's, the stacks degrade and need replacement every few years, and running on natural gas still produces CO2, so the climate case depends on biogas or eventually hydrogen. Hydrogen supply at data center scale does not exist yet in most places.

When to use

Consider fuel cells where air permits block engines, where the utility cannot deliver capacity soon enough and gas is available, and where the site wants a lower-emission story than a gas engine plant. They fit best as continuous prime power with the grid as backup. Do not choose them for pure standby duty, where a diesel is a fraction of the cost and starts on demand. And treat hydrogen-fueled designs as a research direction rather than a plan until a supply contract exists.

Key numbers

Electrical efficiency 50–60% on natural gas · modules typically 200–500 kW, stacked to site capacity · very low NOx and particulate emissions, which is the permitting argument · installed cost several times a gas engine per kW · stack replacement every few years is a scheduled operating cost.

Examples

Bloom Energy servers at Equinix, Apple, and several colocation campuses; Microsoft's hydrogen fuel cell demonstration replacing a diesel generator for backup duty; Korean and Japanese installations using fuel cells for both power and heat.

Videos
How A Bloom Energy Server WorksBloom Energy · 100k+ views
Bloom Energy Explained | Can It Power the AI Boom?Leo Cui, Ph.D., CFA · 10k+ views
From Fuel Cell to Energy Server Farm | Bloom EnergyBloom Energy · 1k+ views
Further reading

Key Questions on Energy and AI (IEA)

Class III

Air cooling

moving heat with air, and its limits5 systems

A computer room air conditioner is a self-contained unit with its own refrigeration circuit: it cools room air with a direct-expansion coil and rejects heat to an outdoor condenser. A computer room air handler has no refrigeration of its own; it is a coil and a fan fed with chilled water from a central plant. Both blow cold air, traditionally into a raised-floor plenum and up through perforated tiles in the cold aisle. The distinction matters because it decides where the compressor lives, and therefore how efficiently the whole site can run.

Strengths & weaknesses

CRAC units are simple and self-contained, which suits small rooms with no chilled water and edge sites with no plant. CRAH units are more efficient at scale, because a central chiller plant with economization beats many small compressors. Both share the same ceiling: a raised floor and perforated tiles can deliver roughly 5–15 kW per rack before airflow, not cooling capacity, becomes the limit. Bypass air and recirculation waste a large share of the fan energy in most legacy rooms, and the fix is containment rather than more units.

When to use

Use CRAH units with a central chilled water plant for any facility above a few hundred kilowatts, since that is where economizers and efficient chillers pay. Use CRAC units for small rooms, edge cabinets, and retrofits where running chilled water piping is impractical. In an existing room, fix air management before adding units: blanking panels, sealed floor cutouts, and correctly placed tiles routinely recover more capacity than another CRAH would. Above about 20 kW per rack, plan the move to liquid rather than adding air capacity you cannot use.

Key numbers

Practical limit of raised-floor air delivery is roughly 5–15 kW per rack · fans are typically 10–20% of cooling energy, and variable speed cuts that sharply · CRAH supply air commonly 18–27 °C under current ASHRAE guidance, up from 13 °C historically · bypass air and recirculation waste a large fraction of airflow in uncontained rooms · CRAC compressors are less efficient at scale than a central chiller plant.

Examples

Raised-floor rooms in nearly every enterprise facility built between 1990 and 2015; CRAH-plus-chiller designs in colocation halls; ASHRAE's successive widening of the recommended inlet envelope, which allowed most sites to raise supply temperature and save compressor energy.

Videos
Computer Room Air Conditioning - How do CRAC units work?The Engineering Mindset · 100k+ views
CRAC vs CRAH Units ExplainedMEP Academy · 10k+ views
The Crucial Role of the CRAH in a Data CenterCoreSite · 10k+ views
Further reading

ASHRAE Data Center Resources, Datacom series (ASHRAE) · Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP)

Containment puts a physical barrier between the cold air going into the servers and the hot air coming out. Cold aisle containment encloses the cold aisle with doors and a roof, so the rest of the room becomes a hot return plenum. Hot aisle containment does the reverse, ducting the hot aisle back to the cooling units and leaving the room cold. Add blanking panels in empty rack units and brushes in floor cutouts and the two air streams stop mixing, which is the entire point.

Strengths & weaknesses

It is the cheapest large efficiency gain available in an air-cooled room. Because supply and return no longer mix, supply temperature can rise several degrees, the delta across the coil widens, fans slow down, and chillers spend more hours on economizer. Sites routinely cut cooling energy 20–40% and gain rack capacity they already paid for. The costs are modest and mostly practical: hot aisle containment makes the contained aisle genuinely hot to work in, fire suppression and sprinkler coverage have to be reviewed, and containment reduces the thermal buffer, so a cooling failure raises inlet temperatures within a minute rather than several.

When to use

Contain every air-cooled hall. There is essentially no case against it in a new build, and in a retrofit it is usually the first thing to do, before adding cooling units or raising set points. Choose hot aisle containment where the room will be occupied and staff comfort matters, and cold aisle containment where retrofitting is easier because the existing units already feed a raised floor. Review fire suppression with the containment in place, and model what happens to inlet temperature during a cooling outage before relying on ride-through assumptions written for an uncontained room.

Key numbers

Cooling energy savings commonly 20–40% in a previously uncontained room · lets supply temperature rise several degrees, which extends economizer hours substantially · payback often under two years · thermal ride-through shrinks to roughly a minute after a cooling failure · blanking panels and sealed cutouts deliver much of the benefit for very little money.

Examples

Universal in hyperscale design since the early 2010s; colocation retrofits where containment released stranded capacity without new mechanical plant; Berkeley Lab's data center best-practice guidance, which puts air management ahead of equipment upgrades.

Videos
Hot Aisle vs Cold Aisle Containment ExplainedMEP Academy · 5k+ views
Data Center Cooling - how are data centre cooled cold aisle containment hvacrThe Engineering Mindset · 100k+ views
Hot and cold aisle in data center explained in simple termsNETWORKING WITH H · 10k+ views
Further reading

Data Center Energy Efficiency Toolkit (Berkeley Lab Center of Expertise)

A central chilled water plant makes cold water in one place and pumps it everywhere it is needed. Chillers, usually water-cooled centrifugal machines for large sites, produce water at 7–18 °C, primary and secondary pumps circulate it, and the load is whatever needs it: air handlers, rear-door heat exchangers, or the facility side of a coolant distribution unit. Heat leaves through cooling towers or dry coolers. This is the backbone that every other cooling technology on this sheet either connects to or deliberately avoids.

Strengths & weaknesses

Central plants are efficient at scale, they can economize when the weather allows, and one water loop serves air cooling and liquid cooling at once, which matters during a mixed transition. Large chillers reach efficiencies small direct-expansion units cannot. The costs are capital, complexity, and water. A plant is a large building commitment with pumps, piping, and controls that all need commissioning, and water-cooled machines consume evaporative water. Raising chilled water temperature is the single most useful efficiency lever available, and most legacy plants run colder than they need to because the set point was chosen for equipment that is long gone.

When to use

Build a central chilled water plant for anything above a few megawatts, and design it for the highest supply temperature the load will accept, since every degree buys economizer hours. Where liquid cooling is coming, size and pipe the plant for higher-temperature loops now, because retrofitting a warm-water circuit later costs far more than allowing for it. Use packaged direct-expansion units instead only at small sites, at the edge, and where site constraints rule out a plant. Whatever the choice, meter the plant properly; most chilled water systems have no idea how much of their output is being wasted.

Key numbers

Chilled water typically supplied at 7–18 °C, and liquid-cooled loads accept far warmer · large centrifugal chillers reach roughly 0.5 kW per ton at design and much better at part load · each degree of higher supply temperature adds economizer hours · water-cooled plants consume evaporative water in the cooling towers · plant and distribution are a large share of the mechanical capital cost.

Examples

Chilled water plants in essentially every large colocation campus; warm-water loops at 32 °C and above in HPC facilities, which allow year-round economization; ASHRAE's water temperature classes, which set the vocabulary for how warm a liquid loop may run.

Videos
Data Center Chilled Water Systems ExplainedMEP Academy · 10k+ views
Chilled Water Central Plant BasicsMEP Academy · 100k+ views
Air Cooled vs WaterCooled Data CentersMEP Academy · 5k+ views
Further reading

Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP)

Economization means using the outdoor environment to do the cooling instead of running a compressor. Air-side economization draws filtered outside air into the hall directly whenever it is cool enough, and exhausts the hot air rather than recirculating it. Water-side economization keeps the building sealed and instead bypasses the chiller when the cooling tower or dry cooler can make cold enough water on its own, usually through a plate heat exchanger. Both trade capital and controls complexity for compressor hours.

Strengths & weaknesses

Compressors are the single largest mechanical load in a data center, and in a cool climate economization can eliminate most of their run hours. Sites in the Pacific Northwest, the Nordics, and northern Europe run essentially chillerless for much of the year, and this is most of the reason those regions attract capacity. The weaknesses split by type. Air-side brings the outdoors inside, so filtration, humidity control, and contamination all need attention, and a smoke event or a chemical release means shutting the dampers. Water-side avoids that but achieves fewer free hours because it needs a bigger temperature difference to work.

When to use

Design for economization anywhere the climate offers meaningful hours, which is most of the temperate world once supply temperature is raised. Prefer water-side where air quality, coastal salt, or humidity control argue against bringing outside air in; prefer air-side where the climate is dry and clean and the extra hours are worth the filtration. The prerequisite for both is a warm supply temperature and good air management, so contain the aisles and raise the set point first. Economization added to a room running 13 °C supply air gets a fraction of the benefit it would at 24 °C.

Key numbers

Compressors are typically the largest mechanical load, so free-cooling hours translate directly into PUE · northern climates achieve several thousand economizer hours a year, and some sites run chillerless · air-side needs filtration and humidity control, plus a shutdown plan for outdoor air events · water-side needs a larger approach temperature and delivers fewer hours · benefits scale with how warm the supply temperature is allowed to run.

Examples

Facebook's Prineville, Oregon facility, which popularized air-side economization at hyperscale; Nordic sites running with almost no compressor hours; water-side economizers retrofitted into existing chilled water plants as the cheapest available efficiency project.

Videos
Air Side EconomizerMEP Academy · 10k+ views
How Waterside Economizers WorkMEP Academy · 10k+ views
Free Cooling - How Does It Work?ICS Cool Energy Ltd · 50k+ views
Further reading

ASHRAE Data Center Resources, Datacom series (ASHRAE)

Evaporating water absorbs a large amount of heat, and evaporative cooling uses that directly. Direct evaporative systems pass outside air through a wetted medium, cooling it toward the wet-bulb temperature before it enters the hall. Indirect systems evaporate water on one side of a heat exchanger and keep the data center's air on the other, so humidity stays controlled. Adiabatic pre-cooling is a lighter version, spraying a mist onto a dry cooler's coil only on the hottest days to extend its capacity.

Strengths & weaknesses

It is very efficient in electricity terms: a fan and a pump replace a compressor, and in a dry climate it can hold supply temperature through summer with no mechanical cooling at all. Adiabatic assist lets a dry cooler be sized for average conditions rather than the design day, which saves capital. The cost is water, and water is now the more visible number. A large evaporative site can consume millions of gallons a year, water quality drives blowdown and treatment, and in drought-prone regions this is a permitting and community issue as much as an engineering one. Direct systems also add humidity that has to be managed.

When to use

Use evaporative cooling in hot dry climates where the wet-bulb temperature is low and water is available and acceptable. Use adiabatic assist widely, since spraying only on peak days captures most of the capital benefit for a small fraction of the water. Avoid direct evaporative in humid climates, where the wet bulb is close to the dry bulb and the technique does little. And check the local water politics before designing around it; several projects have had to redesign to closed-loop cooling after the water number became public.

Key numbers

Approaches the wet-bulb temperature rather than the dry-bulb, so it works best in dry air · replaces compressor power with a fan and a pump · a large evaporative site consumes on the order of millions of gallons a year · adiabatic assist runs only on peak days, so annual water use is a small fraction of full evaporative · water treatment and blowdown are ongoing operating costs.

Examples

Hyperscale sites across Arizona, Nevada, and Spain built around evaporative cooling; Microsoft's shift toward closed-loop designs in water-stressed regions after public scrutiny; adiabatic dry coolers used widely in Europe to extend capacity on a handful of hot days.

Videos
Transtherm Adiabatic Coolers Basics & How it worksTranstherm Cooling Industries Limited · 50k+ views
Direct Evaporative Cooling: How it worksSeeley International EMENA · 50k+ views
Data centers seek sustainable solutions to rising water consumptionCNBC Television · 10k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab)

Class IV

Liquid cooling

cold plates, immersion, and the loops behind them6 systems

A rear-door heat exchanger replaces the rack's back door with a water-cooled coil. Server fans push hot exhaust straight through it, the water takes the heat away, and air leaves the rack at roughly room temperature. Passive versions rely on the servers' own fans; active versions add their own fans to reduce the back pressure and handle more load. From the room's point of view the rack produces no heat at all, which means a hall can gain density without touching its air handling.

Strengths & weaknesses

It is the least invasive liquid cooling there is: no change to the servers, no new coolant inside the IT equipment, no vendor lock, and installation rack by rack in a live hall. It handles 20–60 kW per rack, which covers most non-training workloads. The limits are the water and the ceiling. Water now has to be piped to every rack, with the leak detection and drip management that implies, and door coils use relatively cool water, so they do less to enable warm-water economization than direct-to-chip does. Above roughly 60 kW the door runs out and the heat has to be taken at the chip.

When to use

Rear-door exchangers are the right answer for adding density to an existing air-cooled hall, particularly in colocation where the operator cannot dictate what hardware a tenant installs. They are also a good transition step: pipe the hall for water once, start with doors, and move to direct-to-chip in the same rows later. Do not choose them for a rack above 60 kW or so, and do not expect them to deliver the water temperatures that make chiller-free operation possible. For those, go to cold plates.

Key numbers

Handles roughly 20–60 kW per rack, passive at the lower end and active at the upper · requires no change to servers, so it works with any hardware · uses relatively cool water, typically chilled water rather than a warm loop · installs rack by rack in a live hall · needs leak detection and drip containment at every rack.

Examples

Widely deployed in colocation halls upgrading density without rebuilding; university and research clusters using passive doors on standard servers; vendors including nVent, CoolIT, Motivair, and Vertiv shipping both passive and active designs.

Videos
Rear Door Data Centre CoolingAqua Cooling · 10k+ views
RDHX PRO - Rear Door CoolernVent SCHROFF · 10k+ views
CoolIT Rear Door Heat Exchangers (RDHx)CoolIT Systems · 5k+ views
Further reading

Emergence and Expansion of Liquid Cooling in Mainstream Data Centers (ASHRAE Technical Committee 9.9)

Direct-to-chip cooling puts a cold plate on the components that make the most heat, usually the GPUs and CPUs, and pumps water or a water-glycol mix through it. The coolant stays liquid throughout, which is what "single-phase" means. Because the plate sits directly on the die package, the thermal path is short and the coolant can be much warmer than air would have to be: 30–45 °C supply is normal, which is warm enough that a dry cooler can reject the heat without a chiller for most of the year. The remaining 10–30% of rack heat, from memory, drives, and power supplies, still leaves as air.

Strengths & weaknesses

It is the mainstream answer for AI racks and the one every large GPU platform now ships with. Warm-water operation removes most compressor hours, cold plates handle chip power that air physically cannot, and server fan power drops sharply. The complications are plumbing and residual air. Every server has quick-disconnect couplings, so service means breaking and remaking wet connections, leak detection has to work at rack level, and the hall still needs an air path for the fraction of heat the plates do not catch. Coolant chemistry and filtration become a maintenance discipline that data center operations teams have not traditionally had.

When to use

This is the default for anything above roughly 60–80 kW per rack, and increasingly for anything running current-generation accelerators, because the hardware arrives configured for it. Design the facility water loop warm, 32 °C or above, so the plant can economize. Retain about 20–30% of the air capacity for the components that stay air-cooled, and do not delete the air handling when converting a hall. If the racks are below 40 kW and the hardware is heterogeneous, rear-door exchangers get most of the benefit with none of the wet connections inside the servers.

Key numbers

Removes roughly 70–90% of rack heat, with the balance still leaving as air · supply water typically 30–45 °C, warm enough for chiller-free rejection much of the year · supports well over 100 kW per rack · server fan power falls sharply, which is a direct IT-side energy saving · every server carries quick-disconnect couplings that are serviced wet.

Examples

NVIDIA GB200 NVL72 racks, which ship liquid-cooled and set the current density benchmark; long-standing HPC deployments at Oak Ridge and LRZ, which proved warm-water operation years earlier; CoolIT, Vertiv, Motivair, and Supermicro cold-plate systems in production AI halls.

Videos
Supermicro SuperMinute: Direct to Chip Liquid Cooling SolutionsSupermicro · 5k+ views
Liquid Cooling in AI Data CenterMEP Academy · 10k+ views
Data Center Liquid Cooling Explained: Direct-to-Chip vs. ImmersionidcWeek · under 1k views
Further reading

Emergence and Expansion of Liquid Cooling in Mainstream Data Centers (ASHRAE Technical Committee 9.9) · Liquid and Immersion Cooling Options for Data Centers (Vertiv)

Two-phase direct-to-chip uses a dielectric fluid that boils inside the cold plate. Because evaporation absorbs far more heat per kilogram than a temperature rise does, the same flow removes several times the heat, and the plate holds a nearly constant temperature across its surface while boiling. The vapor travels to a condenser, gives up its heat, and returns as liquid. Flow can be driven by a pump or, in some designs, by the density difference between vapor and liquid alone.

Strengths & weaknesses

Heat flux capability is the reason to look at it: chip powers past 1,500 W and future packages with several accelerators in one module are where single-phase plates start to struggle, and boiling handles them with a smaller temperature difference. Uniform plate temperature also reduces thermal stress on the package. The problems are fluid and containment. The working fluids are engineered dielectrics, several of which are PFAS compounds now facing restriction in Europe, they are expensive, and a two-phase loop must be sealed against vapor loss in a way a water loop does not. Field experience is thin and mostly vendor-run.

When to use

Watch it, pilot it if the roadmap includes packages beyond what single-phase can hold, and do not build a hall around it yet. The decision hinges on fluid availability more than thermodynamics: a system designed around a fluid that gets restricted is a stranded asset, so ask what the fluid is, what its regulatory status is, and what the replacement path would be. For everything shipping today, single-phase direct-to-chip is the lower-risk answer and reaches the required densities.

Key numbers

Boiling removes several times the heat per unit flow compared with a sensible-heat loop · holds a nearly constant plate temperature across the boiling surface · targets chip powers above roughly 1,500 W where single-phase plates get difficult · fluids are engineered dielectrics, several of them PFAS compounds under regulatory pressure · deployments are pilots rather than fleets.

Examples

ZutaCore and Accelsius two-phase cold plate systems; Advanced Cooling Technologies' 200 kW two-phase coolant distribution unit; Open Compute Project working sessions on two-phase performance metrics and PFAS sustainability, which is where the fluid question is being argued out.

Videos
A Closer Look at Two-Phase Liquid CoolingData Center Richness · 10k+ views
The Future of Data Center Cooling Starts Here | ACT’s 200 kW Two-Phase CDUAdvanced Cooling Technologies Inc. · 5k+ views
Direct-to-chip, Two-Phase Cooling Performance Metrics and PFAS SustainabilityOpen Compute Project · 10k+ views
Further reading

Immersion Cooling in Data Centers: A Comprehensive Review of Benefits, Challenges, and Future Directions (Thermal and Fluids Engineering Conference, via NSF PAR)

Single-phase immersion drops whole servers into a bath of dielectric fluid, usually a synthetic or mineral oil, laid out horizontally in a sealed tank. A pump circulates the fluid past the boards and through a heat exchanger. Every component is cooled, not just the processors, so the servers have no fans at all and no air path is needed. A tank replaces a rack, and the room around it becomes an ordinary industrial space rather than a conditioned hall.

Strengths & weaknesses

It cools everything, tolerates very high density, eliminates fan power entirely, and works with warm fluid, so heat rejection can be a dry cooler. Because there is no air, dust, humidity, and acoustic noise all disappear, which suits edge locations and dirty environments. The costs are practical rather than thermal. Servers must be modified: fans out, thermal interface materials and some optics changed, and hard drives sealed or replaced. Servicing means lifting a dripping board out of oil, fluid inventory is expensive and heavy, and floor loading for a full tank is well beyond a normal raised floor. Warranty support from server vendors remains uneven.

When to use

Consider single-phase immersion where density is high, the hardware fleet is uniform enough to modify once, and the site has no legacy air infrastructure to preserve, which describes crypto mining and some purpose-built AI and edge deployments. It is also attractive where dust or humidity make air cooling a maintenance problem. Do not choose it for a mixed colocation hall with frequent hardware changes, or where server vendors will not warrant immersed equipment. For most AI deployments direct-to-chip has become the mainstream answer, largely because it does not require modifying the servers.

Key numbers

Supports well over 100 kW per tank · eliminates server fan power entirely, typically 5–10% of IT load · works with warm fluid, so heat rejection needs no chiller in most climates · fluid inventory is expensive and adds substantial floor loading · servers require modification and vendor warranty terms vary.

Examples

Submer and GRC single-phase systems in European and North American deployments; large-scale use in bitcoin mining, which drove much of the early volume; Intel and Supermicro immersion-ready server programs; edge deployments in dusty or humid sites where sealed tanks avoid filtration.

Videos
How Does Immersion Cooling Work? | Single-Phase Immersion Cooling: Climate-Resilient DatacentersSubmer · 1k+ views
Single-Phase Immersion Cooling vs Direct Liquid Cooling (DLC) | How do they compare? | SubmerSubmer · 1k+ views
Immersion Cooling ExplainedMEP Academy · 1k+ views
Further reading

Enough Hot Air: The Role of Immersion Cooling (arXiv)

Two-phase immersion puts servers in a sealed tank of low-boiling-point dielectric fluid. The fluid boils on the hot components at around 50 °C, vapor rises to a condenser coil in the tank lid, condenses, and rains back down. There are no pumps in the primary loop, because the phase change moves the heat by itself, and the boiling point pins component temperature to a narrow band regardless of load. Thermally it is the most capable approach in this sheet.

Strengths & weaknesses

Heat transfer coefficients are the highest available, the tank is nearly silent with no pumps or fans, and temperature control is inherent rather than regulated. The problems are almost entirely about the fluid. The fluorocarbons that boil in the right range are PFAS compounds, 3M announced it would exit PFAS manufacturing by the end of 2025, and European restrictions are advancing, which puts the supply of the enabling material in doubt. The tank must also be sealed against vapor loss, service means opening a vapor space, and fluid cost per tank is high enough that losses matter commercially, not just environmentally.

When to use

Treat two-phase immersion as parked rather than as an option, unless a non-PFAS fluid with the right boiling point and materials compatibility becomes available and supported. The thermal case was always the strongest of any approach; the commercial case now depends on chemistry that is being regulated out. If extreme heat flux is the requirement today, single-phase direct-to-chip reaches the necessary densities with a supply chain that is not in question, and two-phase direct-to-chip is the nearer alternative if the fluid problem is solved.

Key numbers

Fluid boils at around 50 °C, pinning component temperature to a narrow band · highest heat transfer coefficient of any approach here · no pumps or fans in the primary loop · enabling fluids are PFAS compounds, with 3M exiting PFAS manufacture by the end of 2025 and EU restrictions advancing · fluid cost makes vapor loss a commercial as well as an environmental issue.

Examples

Microsoft's two-phase immersion pilot at Quincy, Washington, the best-documented hyperscale trial; Wiwynn and LiquidStack systems; the Open Compute Project's ongoing work on PFAS alternatives, which is where the future of the approach is being decided.

Videos
Two-Phase Immersion Cooling SystemWiwynn · 10k+ views
What is it? Immersion Cooling in 60 secondsGIGABYTE · 100k+ views
Immersion Cooling in 60 SecondsGIGABYTE · 100k+ views
Further reading

Immersion Cooling in Data Centers: A Comprehensive Review of Benefits, Challenges, and Future Directions (Thermal and Fluids Engineering Conference, via NSF PAR)

A coolant distribution unit is the interface between the building's water and the fluid inside the IT equipment. It contains a heat exchanger, pumps, a filter, an expansion vessel, and controls, and it keeps the two loops separate so that the technical cooling loop can be clean, treated, and held at a chosen temperature and pressure while the facility loop does whatever the plant does. Units come in rack-mounted sizes of tens of kilowatts and floor-standing sizes into the megawatts.

Strengths & weaknesses

Separation is the whole value. The IT loop can be run slightly below room pressure so a leak draws air in rather than pushing coolant out, water chemistry can be controlled independently of the plant, and the facility side never sees the servers. Redundant pumps and a filter make the technical loop maintainable without touching IT. The costs are that the unit is one more critical system with pumps that fail, it adds an approach temperature of a few degrees between the loops, and its capacity and redundancy have to be planned per row rather than per building. Filter and fluid maintenance become a scheduled task.

When to use

Any liquid cooling deployment beyond a single rack needs a CDU, so the questions are size and placement. Use in-rack units for a handful of racks or a colocation tenant who cannot alter the facility. Use large floor-standing units to serve a row or a pod where the fleet is uniform, since one big heat exchanger is more efficient and easier to maintain than a dozen small ones. Design for N+1 pumps, since a CDU failure takes out everything downstream of it. And commission the fluid chemistry program at the same time as the hardware, because contamination shows up months later as blocked cold plates.

Key numbers

Rack-mounted units typically 40–100 kW; floor-standing units several hundred kilowatts to over a megawatt · approach temperature of a few degrees between facility and technical loops · technical loop often run at slightly negative pressure so leaks draw air in · N+1 pumps standard, since everything downstream depends on the unit · filtration and fluid chemistry require a scheduled maintenance program.

Examples

Row-scale CDUs feeding NVIDIA GB200 racks; in-rack units in colocation where the tenant brings liquid cooling into an air-cooled hall; Vertiv, CoolIT, Motivair, and Boyd units across current AI deployments.

Videos
Revolutionizing AI & GPU Cooling: The Power of CDUs (Coolant Distribution Units) in Data CentersHVAC TV · 5k+ views
Facility Coolant Distribution Unit Deep DiveDCX LIQUID COOLING SYSTEMS · 1k+ views
Liquid Cooling Technology in Data Centers: How It Supports AI WorkloadsEquinix · 50k+ views
Further reading

Liquid and Immersion Cooling Options for Data Centers (Vertiv)

Class V

Heat rejection & water

the last step to the atmosphere4 systems

A cooling tower rejects heat by evaporating water. Warm condenser water is sprayed over fill material while a fan pulls air through, a small fraction evaporates, and the rest returns several degrees cooler. Because evaporation approaches the wet-bulb temperature rather than the dry-bulb, a tower can make water colder than the outside air, which is exactly what a water-cooled chiller needs to be efficient. Everything else about a cooling tower is a consequence of running an open water system outdoors.

Strengths & weaknesses

It is the most thermally effective heat rejection available and it makes water-cooled chillers, the most efficient large chillers, practical. Capital cost per ton is low. The costs are water and chemistry. Evaporation is consumptive, dissolved solids concentrate and must be flushed out as blowdown, and the open basin needs biocide treatment with Legionella control as a standing obligation. A large campus can consume millions of gallons a year, and that number is now scrutinized publicly. Plume and drift also constrain siting near occupied buildings.

When to use

Use cooling towers where water is available, affordable, and politically uncontested, and where the site is hot enough that the wet-bulb advantage genuinely matters. They remain the right answer in much of the US south and in humid regions where dry coolers would be badly oversized. Switch to dry or hybrid rejection where water is scarce, where the community has raised it, or where the site can accept warmer loop temperatures, which liquid-cooled IT makes possible. That last point is the important one: warm-water direct-to-chip loops can often reject heat dry, which removes the tower entirely.

Key numbers

Approaches the wet-bulb temperature, typically within 3–5 °C · consumes roughly 1.8 liters per kWh of cooling for a typical open tower · blowdown removes concentrated dissolved solids and adds to total water draw · Legionella control and biocide treatment are permanent operating obligations · capital cost per ton is the lowest of the heat rejection options.

Examples

Open towers on nearly every large water-cooled chiller plant; industry water use effectiveness reporting driven by Green Grid metrics; hyperscale operators publishing water figures alongside energy since about 2022, which changed how the trade-off is discussed.

Videos
Cooling Tower Basic OperationChem-Aqua, Inc. · 100k+ views
Data Center Cooling Methods Explained (Air, Liquid & Immersion Cooling)MEP Academy · 50k+ views
How Data Centers Manage Intense Heat: Cooling Systems ExplainedEquinix · 50k+ views
Further reading

ASHRAE Data Center Resources, Datacom series (ASHRAE)

A dry cooler is a radiator: fluid runs through finned tubes, fans push outside air across them, and heat leaves by conduction and convection with no water consumed. Because there is no evaporation, the fluid cannot get colder than the outdoor dry-bulb temperature, and in practice it lands a few degrees above it. That used to make dry rejection impractical for data centers, whose chilled water was too cold. Liquid-cooled IT changed the arithmetic: a direct-to-chip loop that accepts 40 °C water can be rejected dry almost anywhere.

Strengths & weaknesses

Zero water consumption is the point, and with it goes the water treatment, the blowdown, the Legionella program, and the public conversation. A closed loop stays clean, so fouling and chemistry are far simpler. The costs are area and fan power. Dry coolers are physically large for their capacity and need much more airflow, so fan energy is higher and the equipment yard grows; on the hottest days a dry system either derates or needs adiabatic assist. Chiller efficiency also falls when condensing against warm dry air rather than tower water, which is why the design only works well when the load itself accepts warm fluid.

When to use

Choose dry coolers wherever the IT loop can run warm, which now covers most direct-to-chip and immersion deployments, and in any water-stressed or politically sensitive location. Add adiabatic assist to cover the design day rather than sizing the whole plant for it. Stay with evaporative rejection where the load requires genuinely cold water, where the climate is hot and land is expensive, and where water is cheap and uncontroversial. The general rule: raise the loop temperature first, then decide, because loop temperature is what makes dry rejection viable.

Key numbers

No water consumed in normal operation · fluid lands a few degrees above outdoor dry-bulb, against several degrees above wet-bulb for a tower · significantly higher fan power and footprint per kW rejected · adiabatic assist on peak days avoids sizing the whole plant for the design condition · works well when the IT loop accepts 35–45 °C, which liquid cooling allows.

Examples

Microsoft's closed-loop designs announced for new datacenter builds, which eliminate operational water use; Aligned's closed-loop cooling; European sites using dry coolers with adiabatic assist to hold water use near zero for most of the year; HPC facilities running warm-water loops rejected dry.

Videos
What Is Closed-Loop Cooling? | That’s a Great Question.Aligned Data Centers · under 1k views
How closed-loop cooling works in Microsoft's datacentersMicrosoft Datacenters in Your Community · 1k+ views
Datacenter Closed Loop Cooling Explained in 65 SecondsCoolStickFigureGuy · under 1k views
Further reading

Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP)

Any open cooling system needs a water program. Makeup water is filtered and softened, biocide and scale inhibitor are dosed continuously, conductivity is monitored, and blowdown is discharged when dissolved solids concentrate too far. Reuse changes the source: recycled municipal water, industrial effluent, or on-site treated greywater replaces potable supply. Several large operators now build their own treatment plants so they can run on non-potable water that would otherwise be discharged.

Strengths & weaknesses

Reuse is the highest-leverage change available, because it removes the objection that actually gets raised, which is not water use in the abstract but drinking water use. Cycles of concentration are the other lever: running a tower at higher cycles cuts blowdown and total draw significantly for the cost of tighter chemistry. The complications are that lower-quality source water is harder on equipment, needing more treatment and more frequent cleaning, and that an on-site treatment plant is a real facility with operators and permits. Discharge quality is regulated, so blowdown is not simply drained.

When to use

Run a proper treatment program anywhere there is an open loop; the alternative is fouled fill, scaled condensers, and eventually a Legionella incident. Pursue reclaimed water wherever a municipal purple-pipe supply exists, since it is usually cheaper than potable and removes most of the political exposure. Build on-site treatment only at campus scale, where the volumes justify the plant. And report water use effectiveness alongside PUE, because a site that improved its PUE by moving to evaporative cooling has moved a cost rather than removed one.

Key numbers

Cycles of concentration typically 3–6, and raising them cuts blowdown and total draw · water use effectiveness commonly reported in liters per kWh of IT load · Legionella control is a standing regulatory obligation on open systems · reclaimed municipal water is often cheaper than potable and removes most public objection · discharge quality is permitted, so blowdown has its own limits.

Examples

Google's use of reclaimed and industrial water at several US sites; Microsoft's Silicon Valley campus running on recycled water; Digital Realty and Equinix water use effectiveness reporting; municipal purple-pipe agreements now negotiated as part of data center siting deals.

Videos
The Big Data Center Water ProblemAsianometry · 100k+ views
Data Center Water is a DistractionKyle Hill · 100k+ views
How data centers stay cool while reducing water usage demandsWATE 6 On Your Side · under 1k views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab)

A data center converts almost all the electricity it draws into low-grade heat, and normally throws it away. Heat reuse captures it instead and sells it, usually into a district heating network. Air-cooled halls produce 30–40 °C return air, which is too cool to use directly and needs a heat pump to lift it to the 60–80 °C a network wants. Liquid cooling changes this materially: a direct-to-chip return at 45–50 °C needs far less lifting, and some networks now accept it with a small heat pump or none at all.

Strengths & weaknesses

The heat exists whether or not anyone uses it, so the marginal carbon benefit of displacing a gas boiler is real and the revenue, while modest, is genuine. In cities with existing district heating this is a straightforward commercial arrangement. The obstacles are geography and mismatch. A network has to exist within a few kilometers, the operator has to want a supply whose availability depends on someone else's business, and heat demand is seasonal while the data center runs flat. Contracts also constrain the data center's own operation, since it must now keep return temperature within the network's specification.

When to use

Pursue heat reuse where a district network exists nearby and the local regulator or planning authority values it, which describes most of northern Europe. It fits liquid-cooled facilities much better than air-cooled ones, so it belongs in the design conversation when a hall is converting anyway. Do not build a business case around the heat revenue; it is small next to compute revenue. Treat it as a planning and community asset, since in several jurisdictions offering heat has become part of getting permission to build at all.

Key numbers

Nearly all electricity drawn becomes low-grade heat · air-cooled return air at 30–40 °C needs a heat pump to reach the 60–80 °C district networks want · liquid-cooled return at 45–50 °C needs far less lift · viable only within a few kilometers of an existing network · heat demand is seasonal while data center output is constant.

Examples

Google's heat recovery project in Hamina, Finland, feeding a local network; Fortum's Espoo scheme taking heat from Microsoft datacenters; Stockholm Data Parks, which built the model; Danish and Dutch planning rules that now expect heat reuse from new facilities.

Videos
Finland’s Big Idea: Turning Data Center Waste Into HeatBloomberg Television · 100k+ views
Google’s first-ever heat recovery project for neighbourhoods in FinlandGoogle · 10k+ views
Waste heat from data centresFortum · 1k+ views
Further reading

AI: Five charts that put data-centre energy use - and emissions - into context (Carbon Brief)

Class VI

Facility & siting

building type, ownership model, and location5 systems

A hyperscale campus is a purpose-built site of several buildings, each 30–150 MW, sharing a substation, a water supply, and a security perimeter. Design is standardized and repeated: the same hall, the same electrical block, the same mechanical arrangement, built again and again so construction becomes a manufacturing exercise rather than a bespoke project. The operator owns the compute, so the building can be optimized against its own hardware rather than against a generic tenant specification.

Strengths & weaknesses

Repetition drives everything good about it: cost per megawatt falls, construction schedules compress to 12–24 months per building, and operating practice transfers between sites. Because the operator controls the IT, it can run hot aisle containment at aggressive set points, standardize on one cooling architecture, and design power around its own rack specification. The downsides are concentration and inflexibility. A campus is a very large bet on one location's power, water, and politics, and a standardized design that assumed 30 kW racks is expensive to convert when the next generation needs 130 kW.

When to use

This model belongs to operators with enough demand to fill several buildings and enough control over the hardware to design for it. If you are that buyer, the leverage comes from standardizing early and building repeatedly, and from securing power and water before land. If you are not, colocation or a build-to-suit lease gets similar economics without the balance sheet. The genuine risk to plan for is generational: assume the density specification will change during the campus's life, and pipe for liquid cooling even where the first buildings do not need it.

Key numbers

Individual buildings typically 30–150 MW, campuses 100 MW to over 1 GW · construction of 12–24 months per building once the design repeats · standardized design cuts cost per MW substantially against bespoke construction · US data center electricity reached about 176 TWh in 2023 and is projected to roughly double or triple by 2028 · retrofitting a hall designed for air to liquid is a major cost.

Examples

Meta's Prineville and Odense campuses; Microsoft's Boydton and San Antonio sites; the Amazon campus in Indiana built for Anthropic workloads; Google's Council Bluffs campus, one of the largest single sites in the US.

Economic profile

Hyperscale economics come from repetition and from owning the whole stack, so a saved dollar per watt shows up dozens of times. The dominant risk has shifted from construction to inputs: power availability, transformer lead time, and now GPU supply set the schedule, not concrete. That is why operators are signing power purchase agreements and reserving equipment years ahead, and why "powered land" trades at a premium over ordinary industrial land.

Videos
Microsoft reveals its MASSIVE data center (Full Tour)CNET Highlights · 500k+ views
No Nvidia Chips Needed! Amazon’s New AI Data Center For Anthropic Is Truly MassiveCNBC · 1m+ views
The Full Tour: Wisconsin’s First Hyperscale Data Center ConstructionRCNFRD · 1k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · Key Questions on Energy and AI (IEA)

Prefabricated modular means building the data center in a factory and shipping it. Scope varies: skid-mounted power or cooling plant that arrives tested and only needs connecting, all-in-one containerized halls with racks already installed, or full modular buildings assembled from repeated volumetric units. The common thread is moving work from a congested site with a scarce skilled trade to a controlled factory where the same assembly is built repeatedly.

Strengths & weaknesses

Schedule is the main product. Factory work runs in parallel with site work rather than after it, and deployment times fall from 18–24 months to 6–12. Quality is more consistent because commissioning happens on a production line, and the approach suits places where skilled electrical labor is unavailable. Against that, cost per megawatt is usually higher than a well-executed stick-built project, module dimensions are constrained by what a truck can carry, and the design is fixed at order time, so late changes are expensive. Some designs lock the buyer into one vendor's mechanical and electrical ecosystem.

When to use

Choose prefabricated where speed matters more than capital cost, where site labor is scarce, at remote or edge locations, and for capacity that must be added in increments rather than all at once. Prefabricated power and cooling skids inside a conventional building are the most widely useful version, capturing most of the schedule benefit without committing the whole building to a module vendor. Stick-build where the site has good labor, the program is large enough to amortize a bespoke design, and cost per megawatt is the metric being optimized.

Key numbers

Deployment in roughly 6–12 months against 18–24 for conventional construction · module size limited by road transport, typically 12–14 m long · factory commissioning reduces on-site work and rework · cost per MW usually above a well-executed stick-built project · design locked at order, so late changes are expensive.

Examples

Vertiv and Schneider Electric prefabricated power and cooling modules; Aligned's modular builds in Ohio; the containerized data centers that Microsoft and Google used a decade ago and that returned in a different form for AI capacity; edge modules deployed at cell tower sites.

Videos
Prefabricated Modular Data Center Tour | Vertiv™ SmartMod™ MaxVertiv · 5k+ views
The Future of Prefabricated Modular Data Centers | Schneider ElectricSchneider Electric · 10k+ views
A smarter, faster way to build AI-ready data centers | Vertiv™ OneCoreVertiv · 10k+ views
Further reading

Uptime Institute Reports (Uptime Institute)

Colocation is renting space, power, and cooling in someone else's building while owning the servers yourself. Retail colocation sells by the cabinet or the cage, with the operator providing everything up to the rack. Wholesale sells whole halls or buildings, typically 1 MW and up, with the tenant taking on more of the fit-out and operations. Contracts are priced primarily on committed kilowatts rather than on floor area, which tells you what the operator is really selling.

Strengths & weaknesses

It converts capital into operating expense, gives access to carrier-dense interconnection points that would be impossible to replicate, and removes the need to run a critical facility. Deployment takes weeks instead of years. The costs are per-kilowatt price and constraint. Colocation power is more expensive than self-build at scale, halls built for 5–10 kW racks often cannot take modern density, and liquid cooling in a shared hall requires the operator's cooperation on piping, water treatment, and leak response. Long contracts also lock in a density specification that may age badly.

When to use

Colocate when the requirement is under roughly 5–10 MW, when interconnection to many networks matters, or when speed matters more than unit cost. It is also the right way to enter a new geography before committing to a build. Check three things before signing: the hall's actual per-rack power and cooling limit, whether liquid cooling is supported and on what terms, and how power is billed, since metered against committed changes the economics substantially. Above about 10 MW with a stable forecast, build or lease a whole facility instead.

Key numbers

Retail sold by cabinet, wholesale from about 1 MW · priced on committed kW rather than floor area · deployment in weeks against years for a build · legacy halls commonly cap at 5–15 kW per rack, which excludes current AI hardware · liquid cooling support varies widely and is a contract term, not a given.

Examples

Equinix and Digital Realty in the retail and interconnection market; CyrusOne, Vantage, and QTS in wholesale; the Ashburn and Slough interconnection clusters, where colocation exists mainly for who else is in the building; colocation halls now retrofitting rear-door heat exchangers to take AI tenants.

Videos
What is Colocation & How Does It Work?Interxion: A Digital Realty Company · 500k+ views
The 4 Types of Data Centers Explained | Enterprise, Colocation, Hyperscale & EdgeData Center Resources · 5k+ views
Understanding Colocation Data CentersProvision Networks · 1k+ views
Further reading

Uptime Institute Global Data Center Survey 2025 (Uptime Institute)

An edge data center trades scale for proximity. Instead of one large facility a thousand kilometers away, compute sits in a small enclosure near the users: a cabinet in a cell tower compound, a container in a retail parking lot, a room in a regional exchange. Capacity is typically 50 kW to a few megawatts. The purpose is latency, local data handling, and bandwidth cost, since processing video near where it is generated is far cheaper than backhauling it.

Strengths & weaknesses

Latency of a few milliseconds instead of tens is the product, and for real-time control, autonomous systems, and interactive video that difference is the whole application. Local processing also keeps regulated data inside a jurisdiction. The costs are efficiency and operations. Small sites have poor PUE, typically 1.5–2.0, because the fixed overhead of cooling and power conversion does not shrink with load. Nobody is on site, so everything needs remote hands and remote power cycling, physical security is weaker, and managing hundreds of small sites costs more per kilowatt than managing one large one.

When to use

Deploy at the edge when latency or data locality genuinely requires it, and be strict about that test, because most workloads do not. Content caching, industrial control, retail analytics, and telecom network functions are the durable cases. Design for zero site visits: switched PDUs, out-of-band management, and sealed cooling. Where the requirement is only capacity rather than proximity, a regional colocation facility is cheaper, more efficient, and far easier to run.

Key numbers

Typical capacity 50 kW to a few MW · latency in the low single-digit milliseconds against tens for a distant region · PUE commonly 1.5–2.0 because overhead does not scale down · no staff on site, so remote management is mandatory · per-kilowatt operating cost is higher than any other facility type here.

Examples

Cell tower compounds hosting telecom network functions; EdgeConneX and DataBank regional sites; content delivery caches inside internet exchanges; industrial edge cabinets on factory floors running control and vision workloads.

Videos
What are edge data centres and why are they essential for 5G?Business Standard · 10k+ views
Micro Data Centers for Edge ComputingTripp Lite · 1k+ views
Edge data centersCAREL · 5k+ views
Further reading

Data Center Equipment (ENERGY STAR)

Site selection used to weigh land, fiber, tax, and climate. It now starts with one question: how many megawatts can this location deliver and when. Everything else is secondary, because a site with cheap land, good fiber, and no interconnection date is not a site. The rest of the screen covers water availability and permitting, climate for economization, natural hazard exposure, latency to the target users, construction labor, and the local political appetite for a very large industrial load.

Strengths & weaknesses

Getting this right is the highest-leverage decision in the whole project: it fixes power price for decades, decides whether the facility can economize, and sets how much water the design can use. A well-chosen site makes an ordinary design perform well. The difficulty is that the good sites are known and contested. Powered land with an executed interconnection agreement trades at a large premium, several jurisdictions have introduced moratoria or new rules on large loads, and the utility's answer often depends on a transmission upgrade whose schedule nobody controls.

When to use

Screen on power first, water second, and everything else third. Ask the utility for a real capacity date rather than an expression of interest, and evaluate whether a flexible-load or curtailable agreement would move that date. Check climate against the cooling architecture: a site that supports dry cooling year round removes the water question entirely. Talk to the community before the announcement rather than after, since local opposition has become a genuine schedule risk. And be honest about hazards, because insurance and downtime cost more than the land ever will.

Key numbers

Interconnection studies of one to four years, and longer where a transmission upgrade is needed · powered land with an executed agreement carries a substantial premium over raw industrial land · climate determines economizer hours and therefore both PUE and water use · several US states and European jurisdictions have added rules or moratoria on large data center loads · latency to major population centers sets which workloads the site can serve.

Examples

Northern Virginia, where transmission constraint rather than land now limits growth; ERCOT's large-load queue in Texas; Ireland's effective moratorium in the Dublin region; the shift toward Ohio, Georgia, and the Midwest, driven mainly by available substation capacity.

Videos
Telx - Ideal Data Center Site Selection | Schneider ElectricSchneider Electric · under 1k views
Further reading

Key Questions on Energy and AI (IEA)

Class VII

Racks & operations

standards, interconnect, and how it is run5 systems

The 19-inch rack dates from railway signaling equipment and has survived largely by inertia. The Open Compute Project's Open Rack replaced it for hyperscale use: a 21-inch equipment opening in the same 600 mm floor footprint, a shared DC busbar down the back so servers have no individual power supplies or cords, and centralized power shelves and fans. Open Rack v3 raised the busbar to 48 V, added support for far higher rack power, and standardized the mechanical interfaces for liquid cooling manifolds.

Strengths & weaknesses

Centralizing power conversion is more efficient than dozens of small supplies, removes a large number of cables and connectors, and makes rack-level redundancy simpler. The wider opening gives more room for heatsinks and airflow, and blind-mate busbar connections make servers faster to install and remove. The cost is ecosystem: Open Rack equipment is bought from a smaller set of suppliers, it does not fit a standard 19-inch cabinet, and the design is aimed at operators who buy hundreds of racks of one configuration. For a mixed enterprise fleet it is the wrong shape entirely.

When to use

Adopt Open Rack when buying at hyperscale volume with a uniform hardware fleet, since the efficiency and serviceability gains multiply. It is also increasingly the practical choice for dense AI deployments, because current accelerator racks are designed around this form factor and the liquid cooling manifolds assume it. Stay with 19-inch racks for enterprise and colocation, where equipment comes from many vendors and cabinets have to accept whatever arrives. Whatever the choice, confirm floor loading early; a fully populated liquid-cooled rack can exceed 1,500 kg.

Key numbers

21-inch equipment opening within the same 600 mm floor pitch as a standard rack · 48 V DC busbar in Open Rack v3, replacing per-server power supplies · centralized power shelves are more efficient than distributed supplies · a fully populated liquid-cooled rack can exceed 1,500 kg, well beyond typical raised-floor ratings · specifications are published openly rather than licensed.

Examples

Meta, which created the Open Compute Project and runs Open Rack across its fleet; NVIDIA's GB200 NVL72, built on an Open Rack-derived form factor; Microsoft's contributions of its own rack and cooling designs to OCP; the OCP specification library, which is the reference for interfaces and power.

Videos
OCP Gear Explained - What's inside an OCP Rack?Open Compute Project · 5k+ views
Teaser of the new OCP Academy course series on Open Rack ORv3Open Compute Project · under 1k views
Intro to Rack & Power TrackOpen Compute Project · under 1k views
Further reading

Data Center Energy Efficiency Toolkit (Berkeley Lab Center of Expertise)

Inside a data center, almost everything past a few meters is optical. Structured cabling runs single-mode or multimode fiber from rack to row to spine in a fixed hierarchy, and pluggable transceivers at each end convert electrical signals to light and back. Rates have climbed from 10G through 100G and 400G to 800G per port, and AI clusters have made the network a first-order design problem rather than a utility, because training performance depends on how fast thousands of accelerators can exchange gradients.

Strengths & weaknesses

Fiber carries far more bandwidth per strand than copper, over distances copper cannot reach, with no crosstalk and a very long service life; the cabling plant usually outlives several generations of transceiver. Structured design means adding capacity is a patch rather than a pull. The costs are optics and power. Transceivers are a large share of network capital cost, they consume real power at high rates, and at 800G and above the pluggable module's power becomes a rack-level concern, which is what is pushing co-packaged optics. Cleanliness matters more than people expect; a contaminated connector is the most common fault in the plant.

When to use

Design a structured cabling plant once, generously, and treat it as a 15-year asset while transceivers turn over every few years. Use single-mode where reach or future rate is uncertain, since it costs a little more to install and removes the distance ceiling. In AI clusters, design the network topology alongside the compute rather than after it, and check the optics power budget per rack, because at 800G it is no longer negligible. For short intra-rack links, direct-attach copper remains cheaper and lower power.

Key numbers

Port rates now 400G and 800G, with 1.6T in development · single-mode reaches kilometers, multimode tens to hundreds of meters · transceivers are a large share of network capital and consume meaningful rack power at high rates · cabling plant typically lasts 15 years across several transceiver generations · connector contamination is the most common cause of link faults.

Examples

Spine-and-leaf fabrics in every hyperscale facility; InfiniBand and Ethernet fabrics inside AI training clusters, where interconnect bandwidth limits scaling; co-packaged optics programs from Broadcom and NVIDIA aimed at the transceiver power problem; TIA-942 structured cabling practice, which defines the hierarchy most facilities follow.

Videos
Structured Cabling for Large Data Centers: An Inside Look (Ep. 49)CABLExpress · 10k+ views
Fiber Optic Cabling Solutions for Data Centers | FSFS_com · 1k+ views
What is structured cabling in networking? (Structured Data Cabling)NM Cabling Solutions · 50k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab)

A building management system runs the mechanical and electrical plant: chillers, pumps, air handlers, generators, and switchgear, with alarms and set points. Data center infrastructure management sits alongside it and tracks the IT side: what is in each rack, what each circuit draws, what the inlet temperatures are, and how much power and cooling capacity remains where. The two together answer the question every operator has to answer daily, which is whether the next deployment fits.

Strengths & weaknesses

Done well, DCIM turns capacity planning from an argument into a calculation, catches stranded capacity, and prevents the common failure of a hall that is out of power in one row and empty in another. Asset tracking cuts the time to find and service equipment. The weakness is data quality. A DCIM whose asset records drift from reality is worse than none, because people trust it, and keeping records accurate needs process discipline that many organizations do not sustain. Deployments are also frequently oversold: the software is easy to buy and the operational change is what actually delivers the value.

When to use

Deploy DCIM once the facility is big enough that nobody can hold it in their head, roughly above a few hundred racks or wherever multiple teams share capacity. Start with the measurements that drive decisions, which are per-circuit power and per-rack inlet temperature, and expand from there rather than trying to model everything on day one. Insist on automated data collection wherever possible, since manual entry is where accuracy dies. Keep the building management system on a separate, tightly controlled network, because it can operate plant.

Key numbers

Per-circuit power and per-rack inlet temperature are the two measurements that drive most decisions · stranded capacity of 10–30% is common in facilities without good instrumentation · asset record accuracy decays quickly without automated collection · building management systems control plant and therefore sit inside the security perimeter · payback comes from deferred capacity rather than from energy alone.

Examples

Nlyte, Sunbird, and Schneider EcoStruxure IT in enterprise and colocation; hyperscale operators running their own internal tooling rather than commercial DCIM; colocation providers exposing per-cabinet power data to tenants as a product feature.

Videos
Data Center Infrastructure Management (DCIM) ExplainedAnixter · 50k+ views
What is DCIM? - Data Center Infrastructure Management ExplainedFLUIXAI · 10k+ views
Why DCIM Software is a Game ChangerData Center News · 1k+ views
Further reading

Uptime Institute Reports (Uptime Institute)

Uptime Institute's Tier classification describes how much of a facility can fail or be maintained without stopping the IT load. Tier I is a single path with no redundancy. Tier II adds redundant components. Tier III is concurrently maintainable: any element can be taken out of service for work with the load still running. Tier IV is fault tolerant: an unplanned failure of any single element does not affect the load. Commissioning is the separate discipline of proving the design actually behaves that way, in five levels from factory testing through integrated systems testing under simulated failure.

Strengths & weaknesses

The Tier system gives buyers and designers a shared vocabulary, and a certified design is a genuine commercial signal rather than a marketing claim. Level 5 integrated systems testing, where the whole facility is run at load and faults are deliberately introduced, is the single most effective way to find design and construction errors before a tenant does. The costs are capital and misuse. Tier IV can cost 30–50% more than Tier III for redundancy most workloads no longer need, and the label is widely applied loosely, with "Tier III design" claimed where nothing was ever certified.

When to use

Choose the tier from how the application handles failure. Distributed cloud and AI training workloads that tolerate node and even site loss do not need Tier IV; several hyperscale designs are deliberately below Tier III at the facility level because resilience lives in the software. Enterprise systems with no failover still need concurrent maintainability, and Tier III is the usual answer. Whatever the tier, commission properly and insist on integrated systems testing under load, because an untested redundant design is a redundant design on paper only.

Key numbers

Tier III is concurrently maintainable; Tier IV is fault tolerant against any single unplanned failure · Tier IV typically costs 30–50% more than Tier III · commissioning runs in five levels, ending with integrated systems testing under load · most human-error outages trace to procedures rather than equipment · certification applies to the design, the constructed facility, or the operations, and the three are separate.

Examples

Uptime Institute Tier certifications held by colocation providers as a sales credential; hyperscale designs that deliberately reduce facility redundancy because the application fails over between sites; integrated systems testing catching control-sequence errors that no component test would have found.

Videos
Uptime Data Center Tier Levels - The Gold StandardData Center News · 1k+ views
The 5 Levels of Data Center Commissioning (Explained)Five Nines · 5k+ views
CertMike Explains Data Center TiersMike Chapple · 5k+ views
Further reading

Uptime Institute Global Data Center Survey 2025 (Uptime Institute)

Power usage effectiveness is total facility energy divided by IT energy. A PUE of 1.5 means half a watt of overhead for every watt of computing. Water usage effectiveness is the same idea in liters per kilowatt-hour of IT load. Both are ratios, which is their strength and their weakness: they compare a facility against itself over time honestly, and they compare two facilities against each other only if measured the same way, at the same boundary, over the same period.

Strengths & weaknesses

PUE drove a genuine decade of improvement, because it gave operators one number to manage and a clear target. It is easy to measure once the metering exists, and annualized PUE is hard to game. The weaknesses are what it excludes. PUE says nothing about whether the IT is doing useful work, so replacing servers with more efficient ones makes PUE worse while cutting total energy. It ignores water entirely, which is why WUE exists and why a site can improve PUE by moving to evaporative cooling while making its water position worse. Industry average PUE has been flat near 1.5 for six years.

When to use

Measure annualized PUE and WUE at consistent boundaries and use them to track your own facility. Do not use them to rank facilities in different climates, and do not let a PUE target drive a decision that raises total resource use. For anything about compute efficiency, use a work-per-energy measure instead, since PUE deliberately says nothing about it. When comparing vendor claims, ask what was included, over what period, and at what load, because a design-day figure at full load is a different number from an annualized one.

Key numbers

PUE is total facility energy divided by IT energy; WUE is liters of water per kWh of IT energy · industry weighted average PUE was about 1.54 in 2025, roughly unchanged for six years · large hyperscale sites report trailing PUE around 1.1 · a partly loaded facility has a worse PUE than the same facility at full load · annualized measurement is the only fair basis for comparison.

Examples

The Green Grid's original PUE definition and its later standardization in ISO/IEC 30134; Google's published fleet-wide trailing PUE, among the lowest reported; Uptime Institute survey data showing the industry average stalled; European regulations now requiring reporting of both energy and water for large facilities.

Videos
What is PUE? - Data Center EfficiencyFLUIXAI · 5k+ views
What is PUE, Why Important ? Power Usage Effectiveness Explained | Data Center Efficiency SimplifiedCyber Project Manager EN · under 1k views
What is PUE Finallknoxie7 · 1k+ views
Further reading

Power usage effectiveness (Google Data Centers) · Electrical Efficiency Measurement for Data Centers, White Paper 154 (Schneider Electric)

Glossary

Terms that show up in the system explorer and are not obvious from outside the field. Numbers are typical values, not specifications.

TermWhat it means
Adiabatic assistSpraying a fine mist onto a dry cooler's coil so that evaporation boosts its capacity on hot days. It lets the plant be sized for average conditions instead of the design day, and it uses water only for the handful of hours that need it.
Aisle containmentA physical barrier that keeps cold supply air and hot exhaust air from mixing, by enclosing either the cold aisle or the hot one. It typically cuts cooling energy 20–40% in a room that had none, and it is the cheapest large efficiency gain in air cooling.
BlowdownWater deliberately drained from a cooling tower to stop dissolved solids concentrating as evaporation removes pure water. It adds to total water draw and its quality is regulated, so it cannot simply be discharged.
BuswayAn enclosed conductor bar run above the racks, into which tap-off boxes plug anywhere along its length. It turns adding a circuit into a plug-in operation, and its rating is chosen at design time and expensive to change.
CDUA coolant distribution unit: the heat exchanger, pumps, and controls that separate the building's water from the clean treated fluid circulating through cold plates or rear doors. It lets the two loops run at different temperatures, pressures, and chemistries.
Cold plateA metal block with internal channels, clamped directly onto a processor package, through which coolant flows. Because the thermal path is short, the coolant can be much warmer than air would need to be, which is what makes chiller-free operation possible.
CommissioningProving that what was built behaves the way the design intended, in levels from factory testing to integrated systems testing where the whole facility runs at load and faults are deliberately introduced. It is where control-sequence errors get found, and skipping it is how a redundant design turns out not to be.
CRAC and CRAHA computer room air conditioner has its own refrigeration circuit; a computer room air handler is a coil and fan fed with chilled water from a central plant. The distinction decides where the compressor lives and therefore how efficiently the whole site can run.
Dielectric fluidA liquid that does not conduct electricity, so electronics can be immersed in it or it can flow through sealed loops touching live parts. The good ones for two-phase work are fluorocarbons, which is the reason PFAS regulation matters to cooling design.
Double conversionA UPS topology that rectifies incoming AC to DC and inverts it back, so the load is always fed from the inverter and never sees the utility waveform. Transfer time on a utility failure is zero because there is no transfer.
Dry bulb and wet bulbOrdinary air temperature, and the lowest temperature reachable by evaporating water into that air. Evaporative equipment approaches the wet bulb and dry equipment approaches the dry bulb, which is why humid climates suit dry rejection and arid ones suit evaporative.
EconomizationUsing outdoor conditions to do the cooling instead of running a compressor, either by bringing in outside air or by bypassing the chiller when the tower can make cold enough water. Compressors are the largest mechanical load, so free hours translate directly into efficiency.
Immersion coolingSubmerging whole servers in a bath of dielectric fluid so every component is cooled and no fans are needed. Single-phase circulates the fluid past the boards; two-phase lets it boil on the hot parts and condense on a coil above.
HyperscaleAn operator large enough to build standardized facilities repeatedly and to own the compute inside them, so the building can be designed around its own hardware. Individual buildings run 30–150 MW and campuses can exceed a gigawatt.
N+1 and 2NRedundancy notation. N is what the load needs; N+1 adds one spare unit; 2N duplicates the whole system on independent paths. N+1 usually gives concurrent maintainability, and 2N is what survives a failure during maintenance.
Open RackThe Open Compute Project's rack standard: a 21-inch equipment opening in a standard floor footprint, with a shared DC busbar replacing individual server power supplies and cords. It suits operators buying hundreds of identical racks and does not fit a mixed enterprise fleet.
PDUA power distribution unit. In the room it is the transformer and panel that feeds the racks; in the rack it is the metered strip the servers plug into. Per-outlet metering on the rack version is what capacity planning actually runs on.
PFASPer- and polyfluoroalkyl substances, the fluorinated chemistry behind most two-phase cooling fluids. They persist in the environment, 3M announced an exit from PFAS manufacture by the end of 2025, and EU restrictions are advancing, which puts two-phase supply chains in question.
PUEPower usage effectiveness: total facility energy divided by IT energy. A PUE of 1.5 means half a watt of overhead per watt of computing. It compares a facility against itself honestly and compares different facilities only if the boundary and period match.
Quick disconnectThe dripless coupling that lets a liquid-cooled server be removed without draining the loop. It is what makes cold-plate cooling serviceable, and it is also the component that turns a routine swap into a wet operation.
Raised floorA structural floor on pedestals with a plenum beneath, historically used to deliver cold air through perforated tiles and to route cabling. It caps air delivery at roughly 5–15 kW per rack and rarely carries the weight of a populated liquid-cooled rack.
Rear-door heat exchangerA water-cooled coil that replaces a rack's back door, so exhaust air is cooled before it leaves the cabinet. From the room's point of view the rack produces no heat, which adds density without touching the hall's air handling.
Stranded capacityPower or cooling that exists but cannot be used, because it is in the wrong row, behind the wrong breaker, or blocked by a busway rating. It commonly runs 10–30% in facilities without good instrumentation, and finding it is cheaper than building more.
Tier classificationUptime Institute's scheme for how much of a facility can fail or be maintained without stopping the load. Tier III is concurrently maintainable, Tier IV is fault tolerant against any single unplanned failure, and the difference is 30–50% of capital.
Two-phase and single-phaseWhether the coolant boils. Single-phase stays liquid and carries heat by rising in temperature; two-phase evaporates, which absorbs far more heat per unit flow and holds the surface near the boiling point. Two-phase is thermally better and commercially constrained by fluid chemistry.
WUEWater usage effectiveness: liters of water consumed per kilowatt-hour of IT energy. It exists because a site can improve its PUE by switching to evaporative cooling while making its water position considerably worse.

How to choose data center infrastructure

Two questions decide almost everything else: how many kilowatts per rack, and where does the heat go. Density picks the cooling architecture, cooling architecture picks the loop temperature, and loop temperature decides whether the site can reject heat dry or has to evaporate water. Get those three in order and most of the rest follows. Get them out of order and you build a hall that cannot take the hardware you bought.

Density is the fork in the road

Air can carry a bounded amount of heat out of a cabinet. In practice a well-contained air-cooled hall tops out somewhere between 20 and 40 kW per rack, and beyond that you are moving so much air that fan power and acoustics become the problem. AI training racks are already at 80–140 kW and the roadmaps go higher. That is not an incremental change to an air-cooled design; it is a different building.

Under 10 kW/rack
Legacy enterprise. Raised floor and CRAC units work fine. Fix air management before anything else.
10–40 kW/rack
Mainstream cloud. Contained aisles, chilled water, economization. Air still works if it is done properly.
40–100 kW/rack
Rear-door heat exchangers or direct-to-chip. Water reaches the rack. The hall needs piping.
100 kW+/rack
Direct-to-chip is the default. Power distribution, floor loading, and the building all change with it.

The water-versus-electricity trade

Evaporating water is a cheap way to make cold, so evaporative cooling lowers PUE and raises water use. Rejecting heat dry does the reverse. Neither is right in the abstract; the answer depends on the climate, the local politics of water, and, crucially, on how warm the IT loop can run. That last one is under your control: a direct-to-chip loop that accepts 40 °C water can usually be rejected dry, which removes the water question and much of the compressor load at the same time. Raise the loop temperature first, then choose the rejection method.

Engineering factors

FactorWhy it matters
Rack densityDecides cooling architecture, busway rating, floor loading, and often the building. It is the first number to fix and the hardest to change later.
Loop temperatureThe most useful lever in the whole facility. Warmer loops mean more economizer hours, dry rejection, and lower compressor energy.
Air managementContainment, blanking panels, and sealed cutouts routinely recover 20–40% of cooling energy in a legacy room, for very little money.
Floor loadingA populated liquid-cooled rack can exceed 1,500 kg. Many raised floors cannot take it, and this surprises retrofit projects late.
Redundancy targetTier III concurrent maintainability against Tier IV fault tolerance is a 30–50% capital difference. Choose it from how the application fails over, not from habit.
Ride-throughBatteries give 5–15 minutes, flywheels 15–30 seconds. The shorter one is only acceptable if the generators are genuinely reliable and tested.
Fire and codeLithium batteries, containment, and immersion tanks each change the fire strategy. Involve the authority having jurisdiction before design freeze, not after.
Retrofit disruptionSome upgrades go in rack by rack in a live hall; others need the room emptied. That distinction usually matters more than the capital cost.

Economic and schedule factors

FactorWhy it matters
Power availabilityThe binding constraint on nearly every project. An interconnection date is worth more than a land price, which is why powered land trades at a premium.
Equipment lead timeTransformers, switchgear, and generators all run over a year. Procurement belongs at the front of the schedule.
Stranded capacityFacilities routinely have 10–30% of power or cooling unusable because it is in the wrong place. Instrumentation is what finds it.
Density mismatchA hall designed for 10 kW racks cannot host AI hardware. Existing colocation contracts often lock in a specification the tenant has outgrown.
Water politicsA water number that looks fine in an engineering model can stop a project locally. Reclaimed water and closed loops remove the objection.
Build vs colocateBelow roughly 5–10 MW, colocation is usually faster and cheaper all-in. Above it, self-build wins on unit cost if the demand forecast holds.
UtilizationA half-loaded facility has a worse PUE and worse economics than the same facility full. Phasing the build matters as much as sizing it.

Why the industry average PUE stopped improving

The weighted average PUE reported by Uptime Institute has sat near 1.5 for six years, which surprises people who follow hyperscale announcements of 1.1. Both are true. The easy gains, containment, raised set points, variable-speed fans, and economizers, were captured a decade ago at sites that could take them, and what remains is a long tail of legacy rooms where the building itself is the constraint. New capacity is much better than the average, but it is being added to a fleet that is mostly old. When someone quotes a PUE, ask whether it is one new building or an operator's whole estate.

Core takeaway

Design from the rack outward and from the heat rejection backward, and make them meet. The rack density fixes the cooling architecture; the heat rejection method fixes how much water you spend and how many compressor hours you avoid; the loop temperature is the one variable that improves both at once. Everything else on this sheet is a consequence of those three, and the projects that go wrong are almost always the ones that picked a building first.

Key questions for engineering decisions

Key questions for investment and business analysis

Head-to-head: how do you cool this rack

Rack density is the first fork, so this is the first table. Options are ordered by how much heat they can take out of a cabinet, and the right answer is usually the least invasive one that clears the density you actually need at end of life. The tables after it cover backup power, heat rejection, and how to get the facility at all.

ApproachDensityLoop tempRetrofitPick it when
Air with containmentUp to 20–40 kW18–27 °C supply airDrop-inThe racks fit under 30 kW. Containment is the cheapest capacity in the building and should be done before anything else.
Rear-door heat exchanger20–60 kWChilled waterDrop-in, rack by rackYou need more density in an existing hall and cannot dictate what hardware arrives. The standard colocation answer.
Single-phase direct-to-chip60 kW to well past 10030–45 °CHall retrofitCurrent AI hardware, which ships configured for it. Warm water means the plant can economize or reject dry.
Two-phase direct-to-chipExtreme heat fluxBoiling near chip tempNew buildChip powers past what a single-phase plate can hold. Ask about the fluid's PFAS status before committing.
Single-phase immersion100 kW+ per tankWarm fluidNew buildUniform fleet you can modify, no legacy air plant, and a dusty or humid site. Check server warranties first.
Two-phase immersionHighest availableBoils near 50 °CNew buildThermally the best and commercially parked, because the enabling fluids are PFAS compounds being regulated out.

Carrying the load when the grid drops

Backup is two separate questions: what covers the seconds before the generator picks up, and what covers the hours after. Getting the first wrong is a data loss event; getting the second wrong is a long outage.

OptionRide-throughFootprintReplacement cyclePick it when
Double-conversion UPS with lithium5–15 minutesModerate8–10 yearsThe default. Also gives clean power, and a large plant can bid into demand response between outages.
Double-conversion UPS with lead-acid5–15 minutesLarge3–5 yearsAn existing battery room already sized and suppressed for it, with the replacement cycle already funded.
Flywheel or rotary UPS15–30 secondsSmallAbout 20 yearsGenerators are reliable and tested, batteries are a maintenance burden, and the site runs hot.
Diesel generatorHours to daysLarge yardDecadesStandby duty, essentially always. The constraint is the air permit, not the engine.
On-site gas prime powerContinuousPower plantDecadesInterconnection is years away and the compute cannot wait. Costs more per kWh than the grid.
Fuel cellsContinuousModerateStack every few yearsAir permits block engines. Higher capital, much lower local emissions, still burning gas.

Where the heat finally goes

The last step out of the building is a straight trade between electricity and water, and how much room you have to make that trade depends on how warm the loop runs.

MethodWater useReachesFootprintPick it when
Open cooling towerHighWithin 3–5 °C of wet bulbCompactHot climate, water available and uncontroversial, and the load needs genuinely cold water.
Dry coolerNoneA few degrees above dry bulbLargeThe IT loop runs warm, which liquid cooling allows. Removes the water question entirely.
Adiabatic hybridLowBetween the twoLargeYou want dry operation most of the year without sizing the whole plant for the design day.
Air-side economizerLow to noneOutside air directlyLarge ductsCool clean dry climate. Bring a plan for smoke events and outdoor air quality.
Heat reuse to district networkNoneSells the heatPlant plus connectionA network exists within a few kilometers and the planning authority values it. Not a revenue play.

Build, lease, or rent

Below a certain size, running your own critical facility costs more than it saves. The crossover is mostly about how many megawatts you need and how confident the forecast is.

ModelTime to capacityUnit costFlexibilityPick it when
Hyperscale self-build2–4 yearsLowest at scaleTotal control, large commitmentYou need hundreds of MW, control the hardware, and can carry the balance sheet.
Prefabricated modular6–12 monthsHigher per MWIncrements, vendor-tiedSpeed beats capital cost, or site labor is scarce. Power and cooling skids capture most of the benefit.
Wholesale colocation3–12 monthsMiddleWhole halls, long leases1–10 MW with a stable forecast, and you would rather not operate the building.
Retail colocationWeeksHighest per kWCabinet by cabinetInterconnection density matters, or the requirement is small and uncertain.
EdgeWeeksHighest all-inMany small sitesLatency or data locality genuinely requires proximity. Most workloads do not.