Data Center Infrastructure: A Practical Reference

A decade of data center design assumed 5–10 kW racks and air. AI training racks draw 80–140 kW and are heading past 250, which breaks the air assumption, the power assumption, and usually the building. This guide catalogs 35 systems across seven classes, with the rack density each one actually supports, how much water it uses, and whether it can go into a hall that already exists.

35systems
7classes
8families
Common views
DensityRack power density the system can support. Low under 10 kW/rack is legacy enterprise · Medium 10–40 kW covers most cloud racks · High 40–100 kW is where air stops working · Extreme above 100 kW is AI training. Most systems span a band, and the top of that band is what matters.Each entry covers a span of bands, and picking several widens the results.
Used inWhere the system is normally deployed. AI training is called out separately from hyperscale because its density, power ramp rate, and cooling requirements differ enough to change the design.Pick several tags and an entry has to carry all of them, so each one narrows the results.
Water useOn-site water consumption when the system is running. This is the metric communities and regulators ask about, and it trades directly against electricity: evaporative cooling saves power and spends water, closed loops do the reverse. Untagged where the system does not touch the water balance.Each entry sits in exactly one band, so picking several widens the results.
RetrofitHow disruptive the system is to install in a facility that already exists. Drop-in = goes into a live hall rack by rack · Hall retrofit = needs a room taken out of service and reworked · New build = the building has to be designed around it.Each entry sits in exactly one band, so picking several widens the results.
MaturityMature = standard practice with many suppliers · Established = commercial and widely deployed but newer · Emerging = shipping at scale only in the last few years · Early = deployed in pilots and single sites.Each entry sits in exactly one band, so picking several widens the results.
Class I

Power delivery

from the utility takeoff to the rack outlet4 systems

Everything upstream of the building starts here: a transmission or subtransmission takeoff, an on-site substation with transformers stepping 69–230 kV down to medium voltage, protection, metering, and usually a ring or radial medium-voltage distribution around the campus. A modern hyperscale campus takes 100–1,000 MW, which is a small city's worth of load arriving at one meter, and that is why site selection now begins with a conversation about the substation rather than about fiber or land.

Strengths & weaknesses

Utility service is by far the cheapest power available, at $0.04–0.10/kWh in most US markets, and it comes with no on-site fuel logistics or emissions permit. Reliability from a well-designed dual feed is high. The weaknesses are all schedule and scale. Interconnection studies for a large load run one to four years, the substation transformers behind them run 80–144 weeks, and in constrained regions the answer is simply no until a transmission upgrade lands. Utilities also increasingly ask for curtailment commitments, contribution in aid of construction, or both, in exchange for a faster connection.

When to use

Utility service is the default and should be the plan for any facility that will run for a decade. The decisions are where and on what terms. Choose sites with existing capacity headroom rather than sites that look good on land price, because the interconnection queue dominates the schedule. Where the queue is long, expect to negotiate flexible-load terms, which can move a project years earlier. On-site generation is a bridge or a supplement to this, not a replacement, unless the facility is genuinely remote.

Key numbers

Hyperscale campuses take 100–1,000 MW · service typically at 69–230 kV stepping to 13.8–34.5 kV on site · US industrial power commonly $0.04–0.10/kWh · large-load interconnection studies take one to four years · substation transformer lead times of 80–144 weeks · US data centers used about 176 TWh in 2023, roughly 4.4% of national electricity.

Examples

The northern Virginia cluster, where Dominion's transmission constraints have become the limiting factor on new capacity; ERCOT's large-load interconnection process, built specifically around data centers and crypto; the Ohio and Georgia campuses sited primarily on available substation capacity.

Economic profile

Power cost is 15–30% of a data center's total cost of ownership, so a two-cent difference in rate matters, but availability matters more. Developers now pay premiums for "powered land," a site with an executed interconnection agreement, because that agreement is worth more than the acreage. Where the utility asks for contribution in aid of construction toward network upgrades, that becomes a real capital line item and is worth negotiating alongside the rate.

Videos
Data Center Power Flow: From Utility Grid to Server RackMEP Academy · 50k+ views
Data Center Power Chain - AnimationTechTrainerNJ · 100k+ views
Data Center Power Explained (It's simpler than you think)Base Config · 10k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · Key Questions on Energy and AI (IEA)

Below the UPS, power reaches the racks one of two ways. The traditional route runs cables from a power distribution unit through conduit or under a raised floor to each rack, one circuit at a time. The alternative is overhead busway: a continuous enclosed bar running above the rows, into which tap-off boxes plug anywhere along its length. Busway turns adding a circuit from an electrician's project into a plug-in operation, which is why nearly every hall built for changing IT load now uses it.

Strengths & weaknesses

Busway's advantages are flexibility and airflow. Circuits move as racks move, capacity is added without pulling cable, and getting power out of the underfloor plenum leaves the plenum for air. Metering at the tap-off gives per-rack visibility for free. Against that, busway costs more up front than conduit for a fixed layout, its rating has to be chosen at design time and is expensive to change, and a busway sized for 5 kW racks is now the constraint in many older halls. Cable is cheaper and entirely adequate when the layout will not change.

When to use

Use overhead busway in any hall where rack configurations change, in colocation where tenants come and go, and wherever density is rising, since it is the cheapest part of the chain to over-size. Size it for the density you expect at end of life rather than at day one, because replacing busway means taking the row down. Stay with fixed cable distribution for a small enterprise room with a stable layout, and for retrofits where working overhead is not practical. Whichever route, put revenue-grade metering as far downstream as budget allows; per-rack data is what makes capacity planning possible.

Key numbers

Busway ratings commonly 250–1,200 A, chosen at design and hard to change · tap-off boxes add a circuit in minutes without an outage · per-tap metering is standard on modern systems · overhead routing frees the underfloor plenum for air · cable distribution is cheaper for a fixed layout and worse for a changing one.

Examples

Overhead busway in nearly all colocation halls built since about 2010; underfloor cable distribution in legacy enterprise rooms and in halls that predate high density; busway rating limits now driving hall rebuilds as racks move from 5 kW to 30 kW.

Economic profile

Busway is a small share of the power chain's capital and an expensive part to get wrong. Cable is cheaper for a fixed layout, which is the comparison that wins on day one, but the comparison that matters is what a rating change costs later: a tap-off box adds a circuit in minutes with no outage, while changing the busway itself means taking the row out of service. That asymmetry is why busway is the cheapest part of the chain to over-size, and why ratings chosen for 5 kW racks are now driving hall rebuilds as racks move to 30 kW. The metering is a separate argument. Per-tap data is what locates the 10–30% of installed power that sits stranded in the wrong place, and capacity a site already owns is cheaper than capacity it has to build. In colocation the flexibility is worth more again, because a hall whose circuits move with the tenants can be re-let without an electrician's project each time.

Videos
Is a Busway System Right for Your Data Center?Anixter · 1k+ views
Lesson 7 - Part 2: Power Distribution for Data Centers and UPSEngineering and Donuts · 10k+ views
Further reading

Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP) · Comparing Data Center Power Distribution Architectures, White Paper 129 (Schneider Electric)

A rack power distribution unit is the strip inside the cabinet that turns one or two feeds into the dozens of outlets the servers plug into. Four grades exist: basic, which is a strip; metered, which reports total draw; monitored per-outlet, which reports each outlet; and switched, which can turn outlets on and off remotely. Racks are normally fed from two independent PDUs on separate upstream paths, so either can be lost without dropping dual-corded equipment.

Strengths & weaknesses

The metering is the value. Per-outlet data tells you what is actually drawing power, which is the input to every capacity and stranded-capacity decision, and switched outlets allow a remote power cycle instead of a site visit. They are cheap, install in minutes, and need no design change. The weaknesses are that they are one more thing in a hot cabinet with a finite life, that a failed PDU takes out one power path, and that their built-in circuit breakers trip on inrush if a whole rack is energized at once. At very high density they run out of runway: a 130 kW rack cannot be fed by conventional 208 V strips.

When to use

Use metered or per-outlet PDUs everywhere, since the incremental cost over basic is small and the visibility is what capacity planning runs on. Use switched units in colocation, at the edge, and anywhere staff are not on site. Above about 40–50 kW per rack, stop and look at higher-voltage distribution and rack power shelves instead, because the number of cords and the copper needed to feed the rack conventionally becomes impractical. Always confirm the PDU's own breaker curve against the servers' inrush before a full-rack power-up.

Key numbers

Typical rack feeds 208 V or 415 V three-phase at 30–60 A, giving roughly 5–20 kW per PDU · dual-fed racks are standard, so each path carries the full load in a failure · per-outlet metering typically within plus or minus 1% · switched units allow remote power cycling · conventional strips run out somewhere around 40–50 kW per rack.

Examples

Metered and switched PDUs from Vertiv, Raritan, APC, and ServerTech in nearly every colocation cabinet; per-outlet data feeding DCIM capacity models; 415 V distribution in hyperscale halls, which delivers 240 V line-to-neutral to equipment and cuts conductor size.

Economic profile

Rack PDUs are among the cheapest items in the power chain, and the metering is what a buyer is actually paying for. The step from basic to metered or per-outlet is small, and per-outlet data is the input to every capacity decision in a facility that typically has 10–30% of its installed power stranded in the wrong place. Switched units are justified on a different line: at an edge site or an unstaffed colocation cage, one avoided visit can easily cost more than the unit. The ceiling is what to plan around. Conventional 208 V strips run out somewhere near 40–50 kW per rack, so at AI density the spending moves to rack power shelves and higher-voltage distribution and the PDU stops being the interesting part of the chain. If a hall is being specified for a decade, buy per-outlet metering everywhere and assume the strips themselves get replaced along the way.

Videos
Understanding Rack Power Distribution (PDU)Critical Facilities Connect · under 1k views
Rack Power (PDU) terms and technologyTechTrainerNJ · 100k+ views
Power Distribution Units | Data Center Rack Power Distribution Units | VueNowVueNow Official · 5k+ views
Further reading

Best Practices Guide for Energy-Efficient Data Center Design (Federal Energy Management Program) · Considerations for a Highly Available Intelligent Rack Power Distribution Unit (Vertiv)

Conventional data center power converts several times: AC in, DC in the UPS, AC out, then AC to DC again in every server supply. High-voltage DC distribution cuts the middle out. Power is rectified once at the room or row level and distributed as DC, historically at 380–400 V and now increasingly at plus or minus 400 V DC for AI racks, down a busbar into the rack where power shelves feed the servers. Batteries connect directly to the DC bus, so there is no inverter between them and the load.

Strengths & weaknesses

Removing conversion stages saves a few percent of total facility power and removes hardware that can fail. At AI rack densities the bigger argument is copper: at 130 kW a rack fed at 415 V AC needs an impractical number of large conductors, and raising the voltage is the only way to keep the busbar a reasonable size. Batteries on the DC bus also ride through faster than an inverter can. The weaknesses are ecosystem and safety. DC at 400 V does not self-extinguish an arc the way AC does, so connectors and protection are specialized, the supply chain is thin compared with AC, and equipment has to be bought for it rather than adapted.

When to use

Consider it for new AI training halls at 100 kW or more per rack, where the copper argument alone justifies it and the whole hall can be designed around one architecture. It also fits telecom-adjacent facilities already comfortable with 48 V DC practice. Do not retrofit it into a mixed hall, since running two distribution architectures doubles the spares and the training. For anything under about 50 kW per rack, conventional AC distribution is cheaper, better supported, and efficient enough that the conversion savings do not repay the disruption.

Key numbers

Distribution at 380–400 V DC, and plus or minus 400 V DC in recent AI rack designs · removes two conversion stages, saving roughly 2–5% of facility power · batteries connect directly to the bus with no inverter · DC arcs do not self-extinguish, so protection and connectors are specialized · adopted mainly in new-build AI halls rather than retrofits.

Examples

Open Compute Project rack power designs, which moved from 12 V to 48 V bus bars and then toward higher DC distribution; NVIDIA's 800 V DC reference architecture for high-density AI racks; long-standing 48 V DC practice in telecom central offices, which is the same idea at lower voltage.

Economic profile

The energy saving alone does not pay for this. Removing two conversion stages saves roughly 2–5% of facility power, which is real money at campus scale but not enough to cover specialized DC protection, a thin supply chain, and equipment that has to be bought for the architecture rather than adapted to it. The argument that does pay is copper: at 130 kW a rack fed at 415 V AC needs an impractical number of large conductors, so above about 100 kW the AC alternative stops having an acceptable cost at all. That is why this shows up in new AI halls and almost never as a retrofit, since running two distribution architectures in one building doubles the spares and the training for whatever the second one saves. Volume is what would change the picture. If Open Compute rack power designs and NVIDIA's 800 V DC reference architecture pull enough equipment into the market, the ecosystem premium shrinks and the 2–5% starts to matter on its own. Below about 50 kW per rack, a buyer should stay on AC until that happens.

Videos
Power Distribution for AI Data Centers | Schneider ElectricSchneider Electric · 1k+ views
Further reading

Electrical Efficiency Measurement for Data Centers, White Paper 154 (Schneider Electric) · NVIDIA 800 VDC Architecture Will Power the Next Generation of AI Factories (NVIDIA)

Class II

Standby generation

engines that start when the grid fails2 systems

Standby diesel generators are what carry a data center through a utility outage longer than the batteries can. A 2–3 MW unit in an outdoor enclosure or generator room starts on a signal from the transfer switch, reaches rated speed and voltage in about 10 seconds, and picks up load through an automatic transfer switch. A large campus has dozens of them in N+1 or 2N arrangements, plus fuel storage sized for 24–72 hours of full-load running and a contract for resupply.

Strengths & weaknesses

Diesel generation is the most proven backup there is: fuel is dense and storable, starting is reliable when the maintenance is done, and the units run for years at very low duty. Nothing else combines that availability with on-site fuel. The costs are permitting and testing. Generators are permitted air emission sources with hour limits, and monthly load-bank testing burns fuel and produces the emissions everybody notices. Fuel goes stale and needs polishing, wet stacking damages engines run lightly, and in constrained air-quality districts a large generator plant can be the hardest permit on the project.

When to use

Diesel remains the default for standby duty and should be assumed unless there is a specific reason not to. Sizing follows the redundancy target: N+1 for concurrent maintainability, 2N where the design has to survive a failure during maintenance. Look at alternatives where air permits are the binding constraint, where the site cannot store fuel, or where the utility will pay for the capacity: in some markets, generators enrolled in a demand-response program earn enough to change the business case. Batteries and fuel cells substitute for part of the duty but not for a multi-day outage.

Key numbers

Typical unit 2–3 MW, with dozens per campus · start to full load in about 10 seconds · fuel stored for 24–72 hours at full load · monthly testing under load is standard practice · permitted run hours are usually the binding regulatory constraint · capital roughly $500–900 per kW installed.

Examples

Generator yards at every large colocation and hyperscale campus; Manhattan data centers with generators on the roof and fuel in the basement; Irish and Dutch facilities where generator air permits became a public planning fight; demand-response programs that pay generator fleets to run at peak.

Economic profile

Capital runs roughly $500–900 per kW installed, so a 100 MW site carries $50–90 million of engines that run a few hours a month. The operating cost is mostly testing rather than duty: monthly load-bank runs burn fuel to produce nothing, stored fuel goes stale and needs polishing, and engines run too lightly wet-stack and need repair. None of that is what decides projects. The air permit does, because permitted run hours are the binding constraint, and in a constrained district the generator plant is the hardest permit on the job, which lands as schedule risk rather than as a line in the capital budget. Where a demand-response program pays for the capacity, an operator can earn enough on the same fleet to change the business case, and that is the only common source of revenue from a fleet that is otherwise idle. For standby duty every alternative costs more per kW, so diesel is the sensible default unless the permit blocks it.

Videos
A DAY in the LIFE of the DATA CENTRE | GENERATOR TESTING with ASH!Custodian Data Centres · 100k+ views
What is a Generator and How It Works in a Data Center 1080pCoreSite · 10k+ views
Cat® Diesel Generator Sets Supply Emergency Power to Manhattan’s Data CenterCat Electric Power · 100k+ views
Further reading

Uptime Institute Global Data Center Survey 2025 (Uptime Institute) · Specifics about Provisions Related to Emergency Reciprocating Internal Combustion Engines (US Environmental Protection Agency)

Prime power means the site generates most of its own electricity continuously rather than waiting for an outage. For data centers that currently means reciprocating gas engines or aeroderivative gas turbines behind the meter, sized in the tens to hundreds of megawatts, often with a grid connection retained as backup rather than as the primary supply. The driver is not cost; it is that the interconnection queue in the target market is longer than the business can wait, and gas generation can be permitted and built in 12–24 months.

Strengths & weaknesses

Speed is the product: gas engines arrive as containerized modules and a plant can be running long before a transmission upgrade would finish. The site controls its own capacity and is not exposed to a utility's schedule. Against that, delivered electricity costs more than utility service in most markets once fuel, maintenance, and capital are counted, the site now runs a power plant with the staffing that implies, gas supply needs a pipeline connection with its own lead time, and the air permit for continuous operation is far harder than for standby. Emissions are a live public issue in every jurisdiction where this has been tried.

When to use

Consider prime power where the alternative is waiting years for interconnection and the compute has to be online sooner, and where gas supply and air permits are genuinely obtainable. It also works as a bridge: run on gas now, connect to the grid when the upgrade lands, and keep the plant for backup and peak shaving. Do not choose it where power price drives the model, where the operator has no appetite to run generation, or where the site's emissions profile is a reputational problem. Fuel cells cover a similar niche with lower emissions and higher cost.

Key numbers

Gas engine plants of 10–500 MW, built in 12–24 months against multi-year interconnection waits · reciprocating engines around 40–45% efficient, aeroderivative turbines similar in simple cycle · delivered cost typically above utility rates once fuel and capital are counted · continuous-run air permits are far more restrictive than standby permits · gas pipeline connection carries its own multi-year lead time.

Examples

Texas data centers built with behind-the-meter gas generation while awaiting ERCOT interconnection; xAI's Memphis site, whose gas turbines drew air-permit scrutiny; VoltaGrid and similar modular gas fleets marketed specifically for bridge power; several announced projects pairing gas today with a grid connection later.

Economic profile

Nobody buys this for the power price. Delivered cost sits above utility rates once fuel, maintenance, and capital are counted, but power is only 15–30% of a data center's total cost of ownership, so paying a premium on that share for a few years is a modest penalty against not being able to run at all. What the money buys is 12–24 months to a running plant against a multi-year interconnection wait, so the calculation is whether the operator earns more by starting that much earlier than the fuel premium costs over the life of the plant. The bridge structure improves it: keep the plant for backup and peak shaving after the grid connection lands, and the capital gets a second job instead of being stranded. Two things break the case in diligence. A gas pipeline connection carries its own multi-year lead time, and a continuous-run air permit is far harder to get than a standby permit, so either one can hand back the schedule advantage that justified the plant.

Videos
Inside Climate News: Data Centers Are Building Their Own Gas Power Plants in TexasKXAN · 5k+ views
Powering a Data Center Off the Grid: Part 1 – Natural Gas SolutionsVoltaGrid · under 1k views
VoltaGrid's Data Center SolutionVoltaGrid · 5k+ views
Further reading

Key Questions on Energy and AI (IEA) · AI: Five charts that put data-centre energy use - and emissions - into context (Carbon Brief)

Class II

UPS & ride-through

carrying the load through the gap4 systems

A double-conversion uninterruptible power supply rectifies incoming AC to DC, holds a battery on that DC bus, and inverts back to AC for the load. Because the load is always fed from the inverter, the utility waveform never reaches the IT equipment: sags, harmonics, and frequency wander are all filtered out, and when the utility fails the batteries simply keep supplying the same DC bus with no transfer at all. Ride-through is typically 5–15 minutes, which is far longer than the 10 seconds a generator needs.

Strengths & weaknesses

It is the cleanest power available and the transfer is genuinely seamless. Modern units run at 96–97% efficiency in double conversion, and above 98% in eco mode where the load runs on utility power with the inverter standing by. The costs are capital, footprint, and the batteries. A UPS plant plus batteries is a large room, valve-regulated lead-acid strings need replacing every 3–5 years, and eco mode trades a little transfer risk for the last point of efficiency. Every conversion stage is also a stage that can fail, which is why UPS modules are deployed N+1.

When to use

Double conversion is the default for anything that cannot tolerate a momentary interruption, which is most IT load. Choose it wherever utility power quality is poor, since the filtering is worth as much as the ride-through. Consider eco or multi-mode operation to recover efficiency on a clean supply, with the transfer time verified against the equipment's own tolerance. Where the load is small and the power is clean, a line-interactive unit is cheaper. Where the site wants the batteries to earn money between outages, a lithium battery plant with grid-services capability is the better structure.

Key numbers

Efficiency 96–97% in double conversion, above 98% in eco mode · ride-through typically 5–15 minutes at full load · no transfer time, because the load never leaves the inverter · valve-regulated lead-acid strings last 3–5 years, lithium 8–10 · deployed N+1 at module level in most designs.

Examples

Modular UPS systems from Vertiv, Schneider, Eaton, and ABB in almost every colocation facility; eco-mode operation now common in hyperscale, where power quality is good and the efficiency point is worth chasing; Uptime Institute survey data showing average PUE stuck near 1.5, of which the UPS is a small but persistent contributor.

Economic profile

Most of the cost is the plant and the room it sits in, but the losses run every hour. Modern units hold 96–97% in double conversion and above 98% in eco mode, and one to two points on a 100 MW IT load is 1–2 MW burned continuously, which is why hyperscale operators chase eco mode and sites with poor power quality usually do not. The recurring line item is batteries: valve-regulated lead-acid strings need replacing every 3–5 years against 8–10 for lithium, and each replacement is labor in a live building rather than just cells. Deploying modules N+1 multiplies the whole plant, so the redundancy target moves the capital number more than any efficiency decision does. Where the operator wants the stored energy to do something between outages, a lithium plant with grid-services capability gets two jobs out of the same capital, which is usually the better structure wherever demand response pays.

Videos
How Data Center UPS Systems WorkMEP Academy · 10k+ views
Uninterrupted Power Supply (UPS) Operating modesRockz Automation · 50k+ views
What is a UPS? (Uninterruptible Power Supply)RealPars · 500k+ views
Further reading

Electrical Efficiency Measurement for Data Centers, White Paper 154 (Schneider Electric) · Eaton Energy Saver System: Facts and Principles (Eaton)

A flywheel UPS stores energy as rotation instead of chemistry. A steel or composite rotor spins in a low-friction bearing, usually magnetically levitated in a partial vacuum, and on a power failure the motor becomes a generator and delivers 15–30 seconds of full-load power. That is enough to start a generator, which is all the ride-through most designs need. A diesel rotary UPS goes further and puts the flywheel, a motor-generator, and a diesel engine on one shaft, so the same machine conditions power, rides through, and then runs on fuel.

Strengths & weaknesses

No batteries is the point: nothing to replace on a five-year cycle, no thermal runaway risk, no dedicated battery room with its own cooling and fire suppression, and a 20-year life. Footprint is small for the power and the units tolerate heat that would shorten battery life. Against that, ride-through is seconds rather than minutes, so the design depends completely on the generator starting; the rotating machine needs mechanical maintenance a static UPS does not; and standby losses of 1–2% run continuously. Fewer vendors serve the market than for static systems.

When to use

Choose flywheel or rotary where generators are reliable and tested, where battery replacement cost and room space are real burdens, and where the site is hot enough that batteries would age quickly. Diesel rotary suits large single-block loads and has a strong record in Europe. Do not choose it where the design must survive a generator start failure, since seconds of ride-through leaves no second chance, and do not choose it if the operator wants the stored energy to do anything besides ride-through. Lithium batteries can also provide grid services; a flywheel cannot.

Key numbers

Ride-through 15–30 seconds at full load, against 5–15 minutes for batteries · rotor life around 20 years with no scheduled replacement · standby loss roughly 1–2% of rating · no battery room, no thermal runaway risk · tolerates ambient temperatures that would shorten battery life significantly.

Examples

Active Power and Vycon flywheel systems in North American facilities; Hitec and Piller diesel rotary UPS installations across European data centers; Cisco's Texas facility, an early large rotary deployment; flywheels used as a bridge alongside batteries in hybrid designs.

Economic profile

The case is made over 20 years, not at purchase. A rotor lasts about 20 years with no scheduled replacement, which spans four to six lead-acid replacement cycles at 3–5 years each, and it avoids the battery room with its own cooling and suppression. Floor space in a critical facility is expensive, so the small footprint is part of the return. Against that, standby loss of roughly 1–2% of rating runs continuously, so a comparison that counts only avoided battery spend overstates the saving. Two things narrow the market: fewer vendors serve it than serve static UPS, which shows up in price and in support terms, and the operator gives up the option of bidding stored energy into demand response, which a lithium plant of the same rating allows. A hot site with reliable, well-tested generators is where the numbers come out best, and anywhere else a static UPS with lithium is usually the better buy.

Videos
Data Center World: Flywheel UPS DemonstrationData Center Knowledge · 50k+ views
The Cat Flywheel UPSPeterson Cat · 5k+ views
Dynamic Rotary UPS at Cisco's Allen TX Data Centercycloneinteractive · 50k+ views
Further reading

Reliability Assessment of the Configuration of Dynamic Uninterruptible Power Sources: A Case of Data Centers (Energies) · Comparison of Static and Rotary UPS, White Paper 92 (Schneider Electric)

Lithium-ion has largely displaced valve-regulated lead-acid as the energy store behind the UPS. The cells sit in rack-mounted cabinets with a battery management system that monitors every module, and the chemistry is usually lithium iron phosphate rather than the nickel-rich chemistries used in vehicles, because thermal stability matters more than energy density when the pack lives in a building full of servers. Once the plant is large enough, it stops being only a UPS: the same batteries can shave peaks, respond to grid frequency, or shift load.

Strengths & weaknesses

Compared with lead-acid it takes about a third of the footprint and a quarter of the weight for the same energy, lasts 8–10 years instead of 3–5, tolerates higher ambient temperature, and reports its own state of health. Over a 10-year life the total cost is usually lower despite a higher purchase price. The weaknesses are code and fire. Lithium installations face specific requirements under NFPA 855 and local fire codes, including spacing, detection, and sometimes deflagration venting, and the permitting conversation is materially harder than for lead-acid. Cell supply is also exposed to the same market as vehicles.

When to use

Use lithium for any new UPS energy store where footprint or replacement labor matters, which is nearly all of them, and specify lithium iron phosphate unless there is a specific reason for a nickel chemistry. Size beyond ride-through if the local market pays for demand response or frequency service, because the incremental cells are cheap relative to the rest of the plant. Engage the fire marshal early rather than late. Keep lead-acid where an existing room and its suppression are sized for it and the replacement cycle is already funded.

Key numbers

Roughly a third the footprint and a quarter the weight of lead-acid for the same energy · service life 8–10 years against 3–5 · tolerates higher ambient temperature, so the battery room can run warmer · governed by NFPA 855 and local fire code, which drives spacing and detection · large plants can also bid into demand response and frequency markets.

Examples

Lithium iron phosphate UPS cabinets now standard in new hyperscale builds; Microsoft and Google installations using UPS batteries for grid services; Irish and Dutch facilities offering battery capacity to system operators in exchange for faster connection.

Economic profile

Lithium costs more to buy than lead-acid and usually less over 10 years, because the strings last 8–10 years instead of 3–5 and each replacement is labor in a live building. Footprint is the other half of it: a third the space and a quarter the weight for the same energy frees floor area that could otherwise be white space, which is why the swap is attractive in a hall that already exists and not only in a new build. The cost that surprises people is code. NFPA 855 and the local fire code drive spacing, detection, and sometimes deflagration venting, and the permitting conversation is materially harder than for lead-acid, so bringing the fire marshal in late costs schedule rather than equipment. The upside case is sizing past ride-through: incremental cells are cheap relative to the rest of the plant, so where the local market pays for demand response or frequency service the operator can earn on capacity that would otherwise sit idle. The price risk is that the cells come from the same market as vehicles, so a buyer is bidding against automotive demand.

Videos
What's New with UPS Batteries?Eaton · 1k+ views
ORR Protection Lithium-Ion Battery Q&A: Data Center Code ComplianceORR Protection · under 1k views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · Energy Storage System Guide for Compliance with Safety Codes and Standards (PNNL and Sandia National Laboratories)

A fuel cell converts fuel to electricity electrochemically rather than by combustion. In data centers the dominant type is the solid oxide fuel cell running on natural gas or biogas, delivered as containerized modules of a few hundred kilowatts and stacked to whatever capacity the site needs. Because they run continuously rather than on standby, fuel cells are prime power with a different emissions profile, not a generator substitute. Proton-exchange membrane cells running on hydrogen have been demonstrated in the backup role but remain rare.

Strengths & weaknesses

Electrical efficiency of 50–60% beats a reciprocating engine, and because there is no combustion the NOx and particulate emissions are very low, which is what makes them permittable where engines are not. Modules install quickly and scale in small increments. The costs are capital and fuel. Installed cost per kW runs several times a gas engine's, the stacks degrade and need replacement every few years, and running on natural gas still produces CO2, so the climate case depends on biogas or eventually hydrogen. Hydrogen supply at data center scale does not exist yet in most places.

When to use

Consider fuel cells where air permits block engines, where the utility cannot deliver capacity soon enough and gas is available, and where the site wants a lower-emission story than a gas engine plant. They fit best as continuous prime power with the grid as backup. Do not choose them for pure standby duty, where a diesel is a fraction of the cost and starts on demand. And treat hydrogen-fueled designs as a research direction rather than a plan until a supply contract exists.

Key numbers

Electrical efficiency 50–60% on natural gas · modules typically 200–500 kW, stacked to site capacity · very low NOx and particulate emissions, which is the permitting argument · installed cost several times a gas engine per kW · stack replacement every few years is a scheduled operating cost.

Examples

Bloom Energy servers at Equinix, Apple, and several colocation campuses; Microsoft's hydrogen fuel cell demonstration replacing a diesel generator for backup duty; Korean and Japanese installations using fuel cells for both power and heat.

Economic profile

The capital number does not compete: installed cost per kW runs several times a gas engine's, and stack replacement every few years is a scheduled operating cost an engine does not have. Some of that comes back as fuel, since 50–60% electrical efficiency against 40–45% for a reciprocating engine works out to roughly 10–30% less gas per kWh, but not enough to close the gap on its own. What the premium buys is a permit. Where NOx and particulate limits block engines, the choice is between a fuel cell plant and no site at all, and at that point almost any capital number clears. That is also why fuel cells make no sense for standby duty, since a diesel that runs a few hours a year costs a fraction as much and the emissions argument barely applies at that duty. The climate case depends entirely on the fuel, and biogas costs more than pipeline gas, so an operator buying the low-emission story should price the fuel contract before the hardware.

Videos
How A Bloom Energy Server WorksBloom Energy · 100k+ views
Bloom Energy Explained | Can It Power the AI Boom?Leo Cui, Ph.D., CFA · 10k+ views
From Fuel Cell to Energy Server Farm | Bloom EnergyBloom Energy · 1k+ views
Further reading

Key Questions on Energy and AI (IEA) · Types of Fuel Cells (US Department of Energy)

Class III

Air cooling

moving heat with air, and its limits5 systems

A computer room air conditioner is a self-contained unit with its own refrigeration circuit: it cools room air with a direct-expansion coil and rejects heat to an outdoor condenser. A computer room air handler has no refrigeration of its own; it is a coil and a fan fed with chilled water from a central plant. Both blow cold air, traditionally into a raised-floor plenum and up through perforated tiles in the cold aisle. The distinction matters because it decides where the compressor lives, and therefore how efficiently the whole site can run.

Strengths & weaknesses

CRAC units are simple and self-contained, which suits small rooms with no chilled water and edge sites with no plant. CRAH units are more efficient at scale, because a central chiller plant with economization beats many small compressors. Both share the same ceiling: a raised floor and perforated tiles can deliver roughly 5–15 kW per rack before airflow, not cooling capacity, becomes the limit. Bypass air and recirculation waste a large share of the fan energy in most legacy rooms, and the fix is containment rather than more units.

When to use

Use CRAH units with a central chilled water plant for any facility above a few hundred kilowatts, since that is where economizers and efficient chillers pay. Use CRAC units for small rooms, edge cabinets, and retrofits where running chilled water piping is impractical. In an existing room, fix air management before adding units: blanking panels, sealed floor cutouts, and correctly placed tiles routinely recover more capacity than another CRAH would. Above about 20 kW per rack, plan the move to liquid rather than adding air capacity you cannot use.

Key numbers

Practical limit of raised-floor air delivery is roughly 5–15 kW per rack · fans are typically 10–20% of cooling energy, and variable speed cuts that sharply · CRAH supply air commonly 18–27 °C under current ASHRAE guidance, up from 13 °C historically · bypass air and recirculation waste a large fraction of airflow in uncontained rooms · CRAC compressors are less efficient at scale than a central chiller plant.

Examples

Raised-floor rooms in nearly every enterprise facility built between 1990 and 2015; CRAH-plus-chiller designs in colocation halls; ASHRAE's successive widening of the recommended inlet envelope, which allowed most sites to raise supply temperature and save compressor energy.

Economic profile

The split is capital against energy. A CRAC unit is cheap to buy and needs no plant, so it is the right answer in a small room or an edge cabinet where the alternative is running chilled water to serve a few hundred kilowatts. Above that size a central chiller plant with economization beats many small compressors on operating cost by enough to pay for the piping, which is why almost everything larger is built with CRAH units. In a room that already exists, though, the cheapest capacity is the capacity already installed and being wasted: bypass air and recirculation consume a large fraction of airflow in an uncontained hall, fans are 10–20% of cooling energy, and blanking panels, sealed cutouts, and correctly placed tiles routinely recover more capacity than another CRAH would, for a fraction of the price. The trap is buying air capacity that cannot be delivered. Raised-floor air tops out around 5–15 kW per rack whatever the plant behind it can produce, so past roughly 20 kW per rack an operator adding units is paying for cooling the tiles cannot carry, and the money belongs in liquid instead.

Videos
Computer Room Air Conditioning - How do CRAC units work?The Engineering Mindset · 100k+ views
CRAC vs CRAH Units ExplainedMEP Academy · 10k+ views
The Crucial Role of the CRAH in a Data CenterCoreSite · 10k+ views
Further reading

ASHRAE Data Center Resources, Datacom series (ASHRAE) · Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP)

Containment puts a physical barrier between the cold air going into the servers and the hot air coming out. Cold aisle containment encloses the cold aisle with doors and a roof, so the rest of the room becomes a hot return plenum. Hot aisle containment does the reverse, ducting the hot aisle back to the cooling units and leaving the room cold. Add blanking panels in empty rack units and brushes in floor cutouts and the two air streams stop mixing, which is the entire point.

Strengths & weaknesses

It is the cheapest large efficiency gain available in an air-cooled room. Because supply and return no longer mix, supply temperature can rise several degrees, the delta across the coil widens, fans slow down, and chillers spend more hours on economizer. Sites routinely cut cooling energy 20–40% and gain rack capacity they already paid for. The costs are modest and mostly practical: hot aisle containment makes the contained aisle genuinely hot to work in, fire suppression and sprinkler coverage have to be reviewed, and containment reduces the thermal buffer, so a cooling failure raises inlet temperatures within a minute rather than several.

When to use

Contain every air-cooled hall. There is essentially no case against it in a new build, and in a retrofit it is usually the first thing to do, before adding cooling units or raising set points. Choose hot aisle containment where the room will be occupied and staff comfort matters, and cold aisle containment where retrofitting is easier because the existing units already feed a raised floor. Review fire suppression with the containment in place, and model what happens to inlet temperature during a cooling outage before relying on ride-through assumptions written for an uncontained room.

Key numbers

Cooling energy savings commonly 20–40% in a previously uncontained room · lets supply temperature rise several degrees, which extends economizer hours substantially · payback often under two years · thermal ride-through shrinks to roughly a minute after a cooling failure · blanking panels and sealed cutouts deliver much of the benefit for very little money.

Examples

Universal in hyperscale design since the early 2010s; colocation retrofits where containment released stranded capacity without new mechanical plant; Berkeley Lab's data center best-practice guidance, which puts air management ahead of equipment upgrades.

Economic profile

This is the cheapest large gain available in an air-cooled hall, and it is cheap in two ways: the capital is modest, and the work goes in rack by rack without taking the room out of service. Cooling energy typically falls 20–40% in a previously uncontained room, and payback is often under two years on the energy saving alone. The bigger number is usually capacity. Containment releases racks that were already built and paid for, so a colocation operator can sell more of an existing hall with no new mechanical plant, which is usually a better return than the same money spent on chillers. Blanking panels and sealed floor cutouts deliver much of the benefit for very little, so even the low-budget version pays. What does need budgeting is the review work: fire suppression and sprinkler coverage have to be re-examined with the barriers in place, and thermal ride-through shrinks to roughly a minute, so procedures written for an uncontained room need revisiting too.

Videos
Hot Aisle vs Cold Aisle Containment ExplainedMEP Academy · 5k+ views
Data Center Cooling - how are data centre cooled cold aisle containment hvacrThe Engineering Mindset · 100k+ views
Hot and cold aisle in data center explained in simple termsNETWORKING WITH H · 10k+ views
Further reading

Data Center Airflow Management Retrofit (Federal Energy Management Program) · Implementing Hot and Cold Air Containment in Existing Data Centers, White Paper 153 (Schneider Electric)

A central chilled water plant makes cold water in one place and pumps it everywhere it is needed. Chillers, usually water-cooled centrifugal machines for large sites, produce water at 7–18 °C, primary and secondary pumps circulate it, and the load is whatever needs it: air handlers, rear-door heat exchangers, or the facility side of a coolant distribution unit. Heat leaves through cooling towers or dry coolers. This is the backbone that every other cooling technology on this sheet either connects to or deliberately avoids.

Strengths & weaknesses

Central plants are efficient at scale, they can economize when the weather allows, and one water loop serves air cooling and liquid cooling at once, which matters during a mixed transition. Large chillers reach efficiencies small direct-expansion units cannot. The costs are capital, complexity, and water. A plant is a large building commitment with pumps, piping, and controls that all need commissioning, and water-cooled machines consume evaporative water. Raising chilled water temperature is the single most useful efficiency lever available, and most legacy plants run colder than they need to because the set point was chosen for equipment that is long gone.

When to use

Build a central chilled water plant for anything above a few megawatts, and design it for the highest supply temperature the load will accept, since every degree buys economizer hours. Where liquid cooling is coming, size and pipe the plant for higher-temperature loops now, because retrofitting a warm-water circuit later costs far more than allowing for it. Use packaged direct-expansion units instead only at small sites, at the edge, and where site constraints rule out a plant. Whatever the choice, meter the plant properly; most chilled water systems have no idea how much of their output is being wasted.

Key numbers

Chilled water typically supplied at 7–18 °C, and liquid-cooled loads accept far warmer · large centrifugal chillers reach roughly 0.5 kW per ton at design and much better at part load · each degree of higher supply temperature adds economizer hours · water-cooled plants consume evaporative water in the cooling towers · plant and distribution are a large share of the mechanical capital cost.

Examples

Chilled water plants in essentially every large colocation campus; warm-water loops at 32 °C and above in HPC facilities, which allow year-round economization; ASHRAE's water temperature classes, which set the vocabulary for how warm a liquid loop may run.

Economic profile

Most of the capital sits in the chillers, the pumps, and the distribution piping, and most of the operating cost sits in compressor kilowatt-hours, so the two numbers that move money are how large the plant is and how warm it runs. Large centrifugal machines at roughly 0.5 kW per ton at design, and better than that at part load, are why operators build a central plant above a few megawatts instead of packaged direct-expansion units. Supply temperature is the cheapest lever available, because every degree of higher supply water adds economizer hours and costs nothing to specify at design time. The expensive version of the same decision is making it late: adding a warm-water circuit to a hall piped for 7 °C means taking the room out of service, which is why new plants are increasingly sized and piped for liquid cooling before there is a liquid-cooled rack in the building. Below a few megawatts, or where the building rules out a plant, packaged units are usually cheaper all-in even though they are less efficient.

Videos
Data Center Chilled Water Systems ExplainedMEP Academy · 10k+ views
Chilled Water Central Plant BasicsMEP Academy · 100k+ views
Air Cooled vs WaterCooled Data CentersMEP Academy · 5k+ views
Further reading

Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP) · Chilled Water Plant Design Guide (Energy Design Resources, via Berkeley Lab)

Economization means using the outdoor environment to do the cooling instead of running a compressor. Air-side economization draws filtered outside air into the hall directly whenever it is cool enough, and exhausts the hot air rather than recirculating it. Water-side economization keeps the building sealed and instead bypasses the chiller when the cooling tower or dry cooler can make cold enough water on its own, usually through a plate heat exchanger. Both trade capital and controls complexity for compressor hours.

Strengths & weaknesses

Compressors are the single largest mechanical load in a data center, and in a cool climate economization can eliminate most of their run hours. Sites in the Pacific Northwest, the Nordics, and northern Europe run essentially chillerless for much of the year, and this is most of the reason those regions attract capacity. The weaknesses split by type. Air-side brings the outdoors inside, so filtration, humidity control, and contamination all need attention, and a smoke event or a chemical release means shutting the dampers. Water-side avoids that but achieves fewer free hours because it needs a bigger temperature difference to work.

When to use

Design for economization anywhere the climate offers meaningful hours, which is most of the temperate world once supply temperature is raised. Prefer water-side where air quality, coastal salt, or humidity control argue against bringing outside air in; prefer air-side where the climate is dry and clean and the extra hours are worth the filtration. The prerequisite for both is a warm supply temperature and good air management, so contain the aisles and raise the set point first. Economization added to a room running 13 °C supply air gets a fraction of the benefit it would at 24 °C.

Key numbers

Compressors are typically the largest mechanical load, so free-cooling hours translate directly into PUE · northern climates achieve several thousand economizer hours a year, and some sites run chillerless · air-side needs filtration and humidity control, plus a shutdown plan for outdoor air events · water-side needs a larger approach temperature and delivers fewer hours · benefits scale with how warm the supply temperature is allowed to run.

Examples

Facebook's Prineville, Oregon facility, which popularized air-side economization at hyperscale; Nordic sites running with almost no compressor hours; water-side economizers retrofitted into existing chilled water plants as the cheapest available efficiency project.

Economic profile

Economization is a capital and controls cost bought against compressor kilowatt-hours, so the payback is close to a straight function of how many hours a year the outdoor air can do the work. In a cool climate that is several thousand hours, and some Nordic sites run essentially chillerless, which is most of the reason capacity keeps going to those regions rather than to cheaper land. In a hot humid climate the same equipment buys far fewer hours and often does not pay for itself. The order of spending matters more than the choice between air-side and water-side: containment and a raised set point cost very little, recover 20–40% of cooling energy in a legacy room, and raise the number of economizer hours the next project can claim. An economizer added to a hall still running 13 °C supply air delivers a fraction of the hours it would at 24 °C, so fix the set point first and size the economizer against the fixed room.

Videos
Air Side EconomizerMEP Academy · 10k+ views
How Waterside Economizers WorkMEP Academy · 10k+ views
Free Cooling - How Does It Work?ICS Cool Energy Ltd · 50k+ views
Further reading

ASHRAE Data Center Resources, Datacom series (ASHRAE) · Energy Implications of Economizer Use in California Data Centers (Berkeley Lab)

Evaporating water absorbs a large amount of heat, and evaporative cooling uses that directly. Direct evaporative systems pass outside air through a wetted medium, cooling it toward the wet-bulb temperature before it enters the hall. Indirect systems evaporate water on one side of a heat exchanger and keep the data center's air on the other, so humidity stays controlled. Adiabatic pre-cooling is a lighter version, spraying a mist onto a dry cooler's coil only on the hottest days to extend its capacity.

Strengths & weaknesses

It is very efficient in electricity terms: a fan and a pump replace a compressor, and in a dry climate it can hold supply temperature through summer with no mechanical cooling at all. Adiabatic assist lets a dry cooler be sized for average conditions rather than the design day, which saves capital. The cost is water, and water is now the more visible number. A large evaporative site can consume millions of gallons a year, water quality drives blowdown and treatment, and in drought-prone regions this is a permitting and community issue as much as an engineering one. Direct systems also add humidity that has to be managed.

When to use

Use evaporative cooling in hot dry climates where the wet-bulb temperature is low and water is available and acceptable. Use adiabatic assist widely, since spraying only on peak days captures most of the capital benefit for a small fraction of the water. Avoid direct evaporative in humid climates, where the wet bulb is close to the dry bulb and the technique does little. And check the local water politics before designing around it; several projects have had to redesign to closed-loop cooling after the water number became public.

Key numbers

Approaches the wet-bulb temperature rather than the dry-bulb, so it works best in dry air · replaces compressor power with a fan and a pump · a large evaporative site consumes on the order of millions of gallons a year · adiabatic assist runs only on peak days, so annual water use is a small fraction of full evaporative · water treatment and blowdown are ongoing operating costs.

Examples

Hyperscale sites across Arizona, Nevada, and Spain built around evaporative cooling; Microsoft's shift toward closed-loop designs in water-stressed regions after public scrutiny; adiabatic dry coolers used widely in Europe to extend capacity on a handful of hot days.

Economic profile

The trade is straightforward: a fan and a pump replace a compressor, so the electricity bill falls and the water bill rises. Water is cheap almost everywhere it is available, so the cost that decides projects is the permit rather than the tariff. A large evaporative site consumes on the order of millions of gallons a year, and in a drought-prone region that number can stall an approval or force a late redesign to closed-loop cooling, which is the expensive outcome. Adiabatic assist is the best-value version of the idea: spraying a dry cooler's coil only on the hottest days lets the plant be sized for average conditions rather than the design day, which is a real capital saving, while annual water use stays a small fraction of a fully evaporative site. If you are choosing in a water-stressed market, price the redesign risk alongside the water, because that is the line item that has actually hit projects.

Videos
Transtherm Adiabatic Coolers Basics & How it worksTranstherm Cooling Industries Limited · 50k+ views
Direct Evaporative Cooling: How it worksSeeley International EMENA · 50k+ views
Data centers seek sustainable solutions to rising water consumptionCNBC Television · 10k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · NSIDC Data Center: Energy Reduction Strategies (US Department of Energy Federal Energy Management Program)

Class IV

Liquid cooling

cold plates, immersion, and the loops behind them6 systems

A rear-door heat exchanger replaces the rack's back door with a water-cooled coil. Server fans push hot exhaust straight through it, the water takes the heat away, and air leaves the rack at roughly room temperature. Passive versions rely on the servers' own fans; active versions add their own fans to reduce the back pressure and handle more load. From the room's point of view the rack produces no heat at all, which means a hall can gain density without touching its air handling.

Strengths & weaknesses

It is the least invasive liquid cooling there is: no change to the servers, no new coolant inside the IT equipment, no vendor lock, and installation rack by rack in a live hall. It handles 20–60 kW per rack, which covers most non-training workloads. The limits are the water and the ceiling. Water now has to be piped to every rack, with the leak detection and drip management that implies, and door coils use relatively cool water, so they do less to enable warm-water economization than direct-to-chip does. Above roughly 60 kW the door runs out and the heat has to be taken at the chip.

When to use

Rear-door exchangers are the right answer for adding density to an existing air-cooled hall, particularly in colocation where the operator cannot dictate what hardware a tenant installs. They are also a good transition step: pipe the hall for water once, start with doors, and move to direct-to-chip in the same rows later. Do not choose them for a rack above 60 kW or so, and do not expect them to deliver the water temperatures that make chiller-free operation possible. For those, go to cold plates.

Key numbers

Handles roughly 20–60 kW per rack, passive at the lower end and active at the upper · requires no change to servers, so it works with any hardware · uses relatively cool water, typically chilled water rather than a warm loop · installs rack by rack in a live hall · needs leak detection and drip containment at every rack.

Examples

Widely deployed in colocation halls upgrading density without rebuilding; university and research clusters using passive doors on standard servers; vendors including nVent, CoolIT, Motivair, and Vertiv shipping both passive and active designs.

Economic profile

What a buyer gets from rear doors is avoided disruption rather than better thermal performance. An operator can add density cabinet by cabinet in a hall that is full and under contract, with no room taken out of service and no tenant migration, and in colocation that is worth more than the capital cost of the doors. The offsetting operating cost is that the coils need relatively cool water, so the chillers keep running and the efficiency gain is smaller than a warm direct-to-chip loop would deliver. The other risk is the 60 kW ceiling. An operator who pipes a hall for doors and later has to host 130 kW racks pays for the water distribution once and the cold plates afterwards, which is usually still the cheaper path, because the piping is the part that disrupts a live hall and doing it once is the point of the retrofit.

Videos
Rear Door Data Centre CoolingAqua Cooling · 10k+ views
RDHX PRO - Rear Door CoolernVent SCHROFF · 10k+ views
CoolIT Rear Door Heat Exchangers (RDHx)CoolIT Systems · 5k+ views
Further reading

Emergence and Expansion of Liquid Cooling in Mainstream Data Centers (ASHRAE Technical Committee 9.9) · Data Center Rack Cooling with Rear-door Heat Exchanger (US Department of Energy Federal Energy Management Program)

Direct-to-chip cooling puts a cold plate on the components that make the most heat, usually the GPUs and CPUs, and pumps water or a water-glycol mix through it. The coolant stays liquid throughout, which is what "single-phase" means. Because the plate sits directly on the die package, the thermal path is short and the coolant can be much warmer than air would have to be: 30–45 °C supply is normal, which is warm enough that a dry cooler can reject the heat without a chiller for most of the year. The remaining 10–30% of rack heat, from memory, drives, and power supplies, still leaves as air.

Strengths & weaknesses

It is the mainstream answer for AI racks and the one every large GPU platform now ships with. Warm-water operation removes most compressor hours, cold plates handle chip power that air physically cannot, and server fan power drops sharply. The complications are plumbing and residual air. Every server has quick-disconnect couplings, so service means breaking and remaking wet connections, leak detection has to work at rack level, and the hall still needs an air path for the fraction of heat the plates do not catch. Coolant chemistry and filtration become a maintenance discipline that data center operations teams have not traditionally had.

When to use

This is the default for anything above roughly 60–80 kW per rack, and increasingly for anything running current-generation accelerators, because the hardware arrives configured for it. Design the facility water loop warm, 32 °C or above, so the plant can economize. Retain about 20–30% of the air capacity for the components that stay air-cooled, and do not delete the air handling when converting a hall. If the racks are below 40 kW and the hardware is heterogeneous, rear-door exchangers get most of the benefit with none of the wet connections inside the servers.

Key numbers

Removes roughly 70–90% of rack heat, with the balance still leaving as air · supply water typically 30–45 °C, warm enough for chiller-free rejection much of the year · supports well over 100 kW per rack · server fan power falls sharply, which is a direct IT-side energy saving · every server carries quick-disconnect couplings that are serviced wet.

Examples

NVIDIA GB200 NVL72 racks, which ship liquid-cooled and set the current density benchmark; long-standing HPC deployments at Oak Ridge and LRZ, which proved warm-water operation years earlier; CoolIT, Vertiv, Motivair, and Supermicro cold-plate systems in production AI halls.

Economic profile

The capital is a hall retrofit: piping to every rack, coolant distribution units, manifolds, and enough rework of the plant to supply 30–45 °C. Two operating lines pay for it. Warm supply water lets the site economize or reject heat dry for most of the year, which removes most compressor hours, and server fan power falls sharply, which lands on the IT side of the meter rather than the facility side. What actually decides the purchase is neither of those: current accelerator racks ship configured for cold plates, so anyone buying that hardware is buying the plumbing with it. Budget for the residual air path as well, since 10–30% of rack heat still leaves as air, so a converted hall carries two cooling systems and the air handling cannot be deleted to help pay for the water.

Videos
Supermicro SuperMinute: Direct to Chip Liquid Cooling SolutionsSupermicro · 5k+ views
Liquid Cooling in AI Data CenterMEP Academy · 10k+ views
Data Center Liquid Cooling Explained: Direct-to-Chip vs. ImmersionidcWeek · under 1k views
Further reading

Emergence and Expansion of Liquid Cooling in Mainstream Data Centers (ASHRAE Technical Committee 9.9) · Liquid and Immersion Cooling Options for Data Centers (Vertiv)

Two-phase direct-to-chip uses a dielectric fluid that boils inside the cold plate. Because evaporation absorbs far more heat per kilogram than a temperature rise does, the same flow removes several times the heat, and the plate holds a nearly constant temperature across its surface while boiling. The vapor travels to a condenser, gives up its heat, and returns as liquid. Flow can be driven by a pump or, in some designs, by the density difference between vapor and liquid alone.

Strengths & weaknesses

Heat flux capability is the reason to look at it: chip powers past 1,500 W and future packages with several accelerators in one module are where single-phase plates start to struggle, and boiling handles them with a smaller temperature difference. Uniform plate temperature also reduces thermal stress on the package. The problems are fluid and containment. The working fluids are engineered dielectrics, several of which are PFAS compounds now facing restriction in Europe, they are expensive, and a two-phase loop must be sealed against vapor loss in a way a water loop does not. Field experience is thin and mostly vendor-run.

When to use

Watch it, pilot it if the roadmap includes packages beyond what single-phase can hold, and do not build a hall around it yet. The decision hinges on fluid availability more than thermodynamics: a system designed around a fluid that gets restricted is a stranded asset, so ask what the fluid is, what its regulatory status is, and what the replacement path would be. For everything shipping today, single-phase direct-to-chip is the lower-risk answer and reaches the required densities.

Key numbers

Boiling removes several times the heat per unit flow compared with a sensible-heat loop · holds a nearly constant plate temperature across the boiling surface · targets chip powers above roughly 1,500 W where single-phase plates get difficult · fluids are engineered dielectrics, several of them PFAS compounds under regulatory pressure · deployments are pilots rather than fleets.

Examples

ZutaCore and Accelsius two-phase cold plate systems; Advanced Cooling Technologies' 200 kW two-phase coolant distribution unit; Open Compute Project working sessions on two-phase performance metrics and PFAS sustainability, which is where the fluid question is being argued out.

Economic profile

The economics of two-phase direct-to-chip turn on the fluid rather than on the hardware. Engineered dielectrics cost far more per liter than water or a glycol mix, a sealed loop has to hold vapor rather than merely contain liquid, and anything lost is a purchase rather than a top-up from the tap. Sitting on top of that price is regulatory exposure: several of the candidate fluids are PFAS compounds facing restriction in Europe, so an operator who builds around one of them can end up with equipment that cannot legally be refilled. That is why deployments are pilots rather than fleets, and why the diligence question is the fluid's regulatory status and replacement path rather than its heat flux. Since single-phase cold plates already reach the densities shipping today at a known price, two-phase is worth committing to only if chip packages go past what a single-phase plate can hold and a fluid without the regulatory problem arrives with them.

Videos
A Closer Look at Two-Phase Liquid CoolingData Center Richness · 10k+ views
The Future of Data Center Cooling Starts Here | ACT’s 200 kW Two-Phase CDUAdvanced Cooling Technologies Inc. · 5k+ views
Direct-to-chip, Two-Phase Cooling Performance Metrics and PFAS SustainabilityOpen Compute Project · 10k+ views
Further reading

Immersion Cooling in Data Centers: A Comprehensive Review of Benefits, Challenges, and Future Directions (Thermal and Fluids Engineering Conference, via NSF PAR) · Pumped Two-Phase Learning Center (Advanced Cooling Technologies)

Single-phase immersion drops whole servers into a bath of dielectric fluid, usually a synthetic or mineral oil, laid out horizontally in a sealed tank. A pump circulates the fluid past the boards and through a heat exchanger. Every component is cooled, not just the processors, so the servers have no fans at all and no air path is needed. A tank replaces a rack, and the room around it becomes an ordinary industrial space rather than a conditioned hall.

Strengths & weaknesses

It cools everything, tolerates very high density, eliminates fan power entirely, and works with warm fluid, so heat rejection can be a dry cooler. Because there is no air, dust, humidity, and acoustic noise all disappear, which suits edge locations and dirty environments. The costs are practical rather than thermal. Servers must be modified: fans out, thermal interface materials and some optics changed, and hard drives sealed or replaced. Servicing means lifting a dripping board out of oil, fluid inventory is expensive and heavy, and floor loading for a full tank is well beyond a normal raised floor. Warranty support from server vendors remains uneven.

When to use

Consider single-phase immersion where density is high, the hardware fleet is uniform enough to modify once, and the site has no legacy air infrastructure to preserve, which describes crypto mining and some purpose-built AI and edge deployments. It is also attractive where dust or humidity make air cooling a maintenance problem. Do not choose it for a mixed colocation hall with frequent hardware changes, or where server vendors will not warrant immersed equipment. For most AI deployments direct-to-chip has become the mainstream answer, largely because it does not require modifying the servers.

Key numbers

Supports well over 100 kW per tank · eliminates server fan power entirely, typically 5–10% of IT load · works with warm fluid, so heat rejection needs no chiller in most climates · fluid inventory is expensive and adds substantial floor loading · servers require modification and vendor warranty terms vary.

Examples

Submer and GRC single-phase systems in European and North American deployments; large-scale use in bitcoin mining, which drove much of the early volume; Intel and Supermicro immersion-ready server programs; edge deployments in dusty or humid sites where sealed tanks avoid filtration.

Economic profile

With immersion the spending moves out of the building and into the tank. There is no conditioned hall, no air handling, and no server fans at all, and fan power alone is typically 5–10% of IT load, so the operating case is genuinely strong. Against it sit the fluid inventory, which is expensive and heavy enough to need a floor built for it, and a per-server modification cost paid on every box that goes in and every box swapped out. That is why the economics work best for a uniform fleet that is rarely touched, which describes bitcoin mining, where most of the early volume went. Direct-to-chip took the AI halls for a commercial reason rather than a thermal one: cold plates arrive on a server the vendor already warrants, while immersion asks the buyer to modify the hardware first and then negotiate the warranty terms.

Videos
How Does Immersion Cooling Work? | Single-Phase Immersion Cooling: Climate-Resilient DatacentersSubmer · 1k+ views
Single-Phase Immersion Cooling vs Direct Liquid Cooling (DLC) | How do they compare? | SubmerSubmer · 1k+ views
Immersion Cooling ExplainedMEP Academy · 1k+ views
Further reading

Enough Hot Air: The Role of Immersion Cooling (arXiv) · Data Center Immersion Cooling: A Case Study and Summary of High-Performance Computing Cooling Technologies (Sandia National Laboratories)

Two-phase immersion puts servers in a sealed tank of low-boiling-point dielectric fluid. The fluid boils on the hot components at around 50 °C, vapor rises to a condenser coil in the tank lid, condenses, and rains back down. There are no pumps in the primary loop, because the phase change moves the heat by itself, and the boiling point pins component temperature to a narrow band regardless of load. Thermally it is the most capable approach in this sheet.

Strengths & weaknesses

Heat transfer coefficients are the highest available, the tank is nearly silent with no pumps or fans, and temperature control is inherent rather than regulated. The problems are almost entirely about the fluid. The fluorocarbons that boil in the right range are PFAS compounds, 3M announced it would exit PFAS manufacturing by the end of 2025, and European restrictions are advancing, which puts the supply of the enabling material in doubt. The tank must also be sealed against vapor loss, service means opening a vapor space, and fluid cost per tank is high enough that losses matter commercially, not just environmentally.

When to use

Treat two-phase immersion as parked rather than as an option, unless a non-PFAS fluid with the right boiling point and materials compatibility becomes available and supported. The thermal case was always the strongest of any approach; the commercial case now depends on chemistry that is being regulated out. If extreme heat flux is the requirement today, single-phase direct-to-chip reaches the necessary densities with a supply chain that is not in question, and two-phase direct-to-chip is the nearer alternative if the fluid problem is solved.

Key numbers

Fluid boils at around 50 °C, pinning component temperature to a narrow band · highest heat transfer coefficient of any approach here · no pumps or fans in the primary loop · enabling fluids are PFAS compounds, with 3M exiting PFAS manufacture by the end of 2025 and EU restrictions advancing · fluid cost makes vapor loss a commercial as well as an environmental issue.

Examples

Microsoft's two-phase immersion pilot at Quincy, Washington, the best-documented hyperscale trial; Wiwynn and LiquidStack systems; the Open Compute Project's ongoing work on PFAS alternatives, which is where the future of the approach is being decided.

Economic profile

The blocking cost is a consumable with no secure supply. The fluids that boil in the right range are PFAS compounds, 3M said it would leave PFAS manufacture by the end of 2025, and European restrictions are advancing, so a buyer is being asked to underwrite a facility on a fluid that may not be purchasable for the second half of that facility's life. Fluid cost is high enough that vapor loss is an operating expense, and every service event opens a vapor space. The thermal case is the strongest of any approach on this sheet and it is still hard to finance, because the risk sits in the chemistry rather than in the engineering. Underwrite it only after a non-PFAS fluid with the right boiling point and materials compatibility is in production and supported by the equipment vendors.

Videos
Two-Phase Immersion Cooling SystemWiwynn · 10k+ views
What is it? Immersion Cooling in 60 secondsGIGABYTE · 100k+ views
Immersion Cooling in 60 SecondsGIGABYTE · 100k+ views
Further reading

Immersion Cooling in Data Centers: A Comprehensive Review of Benefits, Challenges, and Future Directions (Thermal and Fluids Engineering Conference, via NSF PAR) · Next Generation Heat Transfer Fluids for Two-Phase Immersion Cooling of Data Centers (Oak Ridge National Laboratory)

A coolant distribution unit is the interface between the building's water and the fluid inside the IT equipment. It contains a heat exchanger, pumps, a filter, an expansion vessel, and controls, and it keeps the two loops separate so that the technical cooling loop can be clean, treated, and held at a chosen temperature and pressure while the facility loop does whatever the plant does. Units come in rack-mounted sizes of tens of kilowatts and floor-standing sizes into the megawatts.

Strengths & weaknesses

Separation is the whole value. The IT loop can be run slightly below room pressure so a leak draws air in rather than pushing coolant out, water chemistry can be controlled independently of the plant, and the facility side never sees the servers. Redundant pumps and a filter make the technical loop maintainable without touching IT. The costs are that the unit is one more critical system with pumps that fail, it adds an approach temperature of a few degrees between the loops, and its capacity and redundancy have to be planned per row rather than per building. Filter and fluid maintenance become a scheduled task.

When to use

Any liquid cooling deployment beyond a single rack needs a CDU, so the questions are size and placement. Use in-rack units for a handful of racks or a colocation tenant who cannot alter the facility. Use large floor-standing units to serve a row or a pod where the fleet is uniform, since one big heat exchanger is more efficient and easier to maintain than a dozen small ones. Design for N+1 pumps, since a CDU failure takes out everything downstream of it. And commission the fluid chemistry program at the same time as the hardware, because contamination shows up months later as blocked cold plates.

Key numbers

Rack-mounted units typically 40–100 kW; floor-standing units several hundred kilowatts to over a megawatt · approach temperature of a few degrees between facility and technical loops · technical loop often run at slightly negative pressure so leaks draw air in · N+1 pumps standard, since everything downstream depends on the unit · filtration and fluid chemistry require a scheduled maintenance program.

Examples

Row-scale CDUs feeding NVIDIA GB200 racks; in-rack units in colocation where the tenant brings liquid cooling into an air-cooled hall; Vertiv, CoolIT, Motivair, and Boyd units across current AI deployments.

Economic profile

A CDU is unavoidable overhead on any liquid deployment beyond a single rack, so the question is what it costs per kilowatt served and across how many racks that cost spreads. Floor-standing units of several hundred kilowatts to over a megawatt cost less per kilowatt than a stack of 40–100 kW in-rack units and are easier to maintain, so a uniform row is the cheaper thing to build. Colocation tenants buy in-rack units anyway at the higher unit cost, because what they are paying for is the ability to bring liquid cooling into a hall they do not own and cannot modify. Two costs get left out of budgets: N+1 pumps, which are not optional when everything downstream depends on one unit, and a fluid chemistry and filtration program, which is a recurring operating cost data center teams have not historically carried. Skipping the second is the expensive mistake, because contamination shows up months later as blocked cold plates in hardware worth far more than the CDU.

Videos
Revolutionizing AI & GPU Cooling: The Power of CDUs (Coolant Distribution Units) in Data CentersHVAC TV · 5k+ views
Facility Coolant Distribution Unit Deep DiveDCX LIQUID COOLING SYSTEMS · 1k+ views
Liquid Cooling Technology in Data Centers: How It Supports AI WorkloadsEquinix · 50k+ views
Further reading

Liquid and Immersion Cooling Options for Data Centers (Vertiv) · Water-Cooled Servers: Common Designs, Components, and Processes (ASHRAE Technical Committee 9.9)

Class V

Heat rejection & water

the last step to the atmosphere4 systems

A cooling tower rejects heat by evaporating water. Warm condenser water is sprayed over fill material while a fan pulls air through, a small fraction evaporates, and the rest returns several degrees cooler. Because evaporation approaches the wet-bulb temperature rather than the dry-bulb, a tower can make water colder than the outside air, which is exactly what a water-cooled chiller needs to be efficient. Everything else about a cooling tower is a consequence of running an open water system outdoors.

Strengths & weaknesses

It is the most thermally effective heat rejection available and it makes water-cooled chillers, the most efficient large chillers, practical. Capital cost per ton is low. The costs are water and chemistry. Evaporation is consumptive, dissolved solids concentrate and must be flushed out as blowdown, and the open basin needs biocide treatment with Legionella control as a standing obligation. A large campus can consume millions of gallons a year, and that number is now scrutinized publicly. Plume and drift also constrain siting near occupied buildings.

When to use

Use cooling towers where water is available, affordable, and politically uncontested, and where the site is hot enough that the wet-bulb advantage genuinely matters. They remain the right answer in much of the US south and in humid regions where dry coolers would be badly oversized. Switch to dry or hybrid rejection where water is scarce, where the community has raised it, or where the site can accept warmer loop temperatures, which liquid-cooled IT makes possible. That last point is the important one: warm-water direct-to-chip loops can often reject heat dry, which removes the tower entirely.

Key numbers

Approaches the wet-bulb temperature, typically within 3–5 °C · consumes roughly 1.8 liters per kWh of cooling for a typical open tower · blowdown removes concentrated dissolved solids and adds to total water draw · Legionella control and biocide treatment are permanent operating obligations · capital cost per ton is the lowest of the heat rejection options.

Examples

Open towers on nearly every large water-cooled chiller plant; industry water use effectiveness reporting driven by Green Grid metrics; hyperscale operators publishing water figures alongside energy since about 2022, which changed how the trade-off is discussed.

Economic profile

Towers are still the cheapest heat rejection to buy, with the lowest capital cost per ton of the options here, and they hold condenser water within 3–5 °C of wet bulb, which is what makes the most efficient large chillers practical. The operating cost has three parts: makeup water at roughly 1.8 liters per kWh of cooling, the chemistry and blowdown an open basin requires, and a Legionella control program that never ends. None of those usually stops a project. The published water total does, now that operators report water alongside energy, so in a drought-prone or contested market the cost that decides a site is a permitting delay rather than a utility bill. If water is cheap and uncontested and the climate is hot and humid, towers are still the right answer on both capital and energy; if not, raising the IT loop temperature enough to reject heat dry removes the tower and the water program with it.

Videos
Cooling Tower Basic OperationChem-Aqua, Inc. · 100k+ views
Data Center Cooling Methods Explained (Air, Liquid & Immersion Cooling)MEP Academy · 50k+ views
How Data Centers Manage Intense Heat: Cooling Systems ExplainedEquinix · 50k+ views
Further reading

ASHRAE Data Center Resources, Datacom series (ASHRAE) · Controlling Legionella in Cooling Towers (US Centers for Disease Control and Prevention)

A dry cooler is a radiator: fluid runs through finned tubes, fans push outside air across them, and heat leaves by conduction and convection with no water consumed. Because there is no evaporation, the fluid cannot get colder than the outdoor dry-bulb temperature, and in practice it lands a few degrees above it. That used to make dry rejection impractical for data centers, whose chilled water was too cold. Liquid-cooled IT changed the arithmetic: a direct-to-chip loop that accepts 40 °C water can be rejected dry almost anywhere.

Strengths & weaknesses

Zero water consumption is the point, and with it goes the water treatment, the blowdown, the Legionella program, and the public conversation. A closed loop stays clean, so fouling and chemistry are far simpler. The costs are area and fan power. Dry coolers are physically large for their capacity and need much more airflow, so fan energy is higher and the equipment yard grows; on the hottest days a dry system either derates or needs adiabatic assist. Chiller efficiency also falls when condensing against warm dry air rather than tower water, which is why the design only works well when the load itself accepts warm fluid.

When to use

Choose dry coolers wherever the IT loop can run warm, which now covers most direct-to-chip and immersion deployments, and in any water-stressed or politically sensitive location. Add adiabatic assist to cover the design day rather than sizing the whole plant for it. Stay with evaporative rejection where the load requires genuinely cold water, where the climate is hot and land is expensive, and where water is cheap and uncontroversial. The general rule: raise the loop temperature first, then decide, because loop temperature is what makes dry rejection viable.

Key numbers

No water consumed in normal operation · fluid lands a few degrees above outdoor dry-bulb, against several degrees above wet-bulb for a tower · significantly higher fan power and footprint per kW rejected · adiabatic assist on peak days avoids sizing the whole plant for the design condition · works well when the IT loop accepts 35–45 °C, which liquid cooling allows.

Examples

Microsoft's closed-loop designs announced for new datacenter builds, which eliminate operational water use; Aligned's closed-loop cooling; European sites using dry coolers with adiabatic assist to hold water use near zero for most of the year; HPC facilities running warm-water loops rejected dry.

Economic profile

Going dry buys out the whole water program: no makeup, no blowdown, no biocide, no Legionella obligation, and no water number to defend at a hearing. The price is land and fan energy, since a dry cooler is physically large for its capacity and moves much more air, and on a hot site it either derates on the design day or needs adiabatic assist to cover it. Whether that trade is a saving depends on how warm the IT loop runs. If the load still needs genuinely cold water, condensing against warm dry air drops chiller efficiency and the site pays in fan power and compressor power at once. If the loop accepts 35–45 °C, which direct-to-chip and immersion allow, the chillers can sit idle for most of the year and dry rejection is cheaper on energy as well as on water, which is why the new builds designed around liquid cooling are the ones going closed-loop.

Videos
What Is Closed-Loop Cooling? | That’s a Great Question.Aligned Data Centers · under 1k views
How closed-loop cooling works in Microsoft's datacentersMicrosoft Datacenters in Your Community · 1k+ views
Datacenter Closed Loop Cooling Explained in 65 SecondsCoolStickFigureGuy · under 1k views
Further reading

Best Practices Guide for Energy-Efficient Data Center Design (Berkeley Lab and FEMP) · Thermosyphon Cooler Hybrid System for Water Savings in an Energy-Efficient HPC Data Center: Results from 24 Months and the Impact on Water Usage Effectiveness (National Laboratory of the Rockies)

Any open cooling system needs a water program. Makeup water is filtered and softened, biocide and scale inhibitor are dosed continuously, conductivity is monitored, and blowdown is discharged when dissolved solids concentrate too far. Reuse changes the source: recycled municipal water, industrial effluent, or on-site treated greywater replaces potable supply. Several large operators now build their own treatment plants so they can run on non-potable water that would otherwise be discharged.

Strengths & weaknesses

Reuse is the highest-leverage change available, because it removes the objection that actually gets raised, which is not water use in the abstract but drinking water use. Cycles of concentration are the other lever: running a tower at higher cycles cuts blowdown and total draw significantly for the cost of tighter chemistry. The complications are that lower-quality source water is harder on equipment, needing more treatment and more frequent cleaning, and that an on-site treatment plant is a real facility with operators and permits. Discharge quality is regulated, so blowdown is not simply drained.

When to use

Run a proper treatment program anywhere there is an open loop; the alternative is fouled fill, scaled condensers, and eventually a Legionella incident. Pursue reclaimed water wherever a municipal purple-pipe supply exists, since it is usually cheaper than potable and removes most of the political exposure. Build on-site treatment only at campus scale, where the volumes justify the plant. And report water use effectiveness alongside PUE, because a site that improved its PUE by moving to evaporative cooling has moved a cost rather than removed one.

Key numbers

Cycles of concentration typically 3–6, and raising them cuts blowdown and total draw · water use effectiveness commonly reported in liters per kWh of IT load · Legionella control is a standing regulatory obligation on open systems · reclaimed municipal water is often cheaper than potable and removes most public objection · discharge quality is permitted, so blowdown has its own limits.

Examples

Google's use of reclaimed and industrial water at several US sites; Microsoft's Silicon Valley campus running on recycled water; Digital Realty and Equinix water use effectiveness reporting; municipal purple-pipe agreements now negotiated as part of data center siting deals.

Economic profile

The chemicals and the labor are a small line item covering a large one, since fouled fill, scaled condensers, and a Legionella incident all cost far more than the program that prevents them. The cheapest improvement available is raising cycles of concentration: moving up within the usual range of 3–6 cuts blowdown and total draw for the price of tighter chemistry and no capital. Reclaimed municipal water is usually cheaper per gallon than potable, and it also removes the objection that actually gets raised, which is drinking water use rather than water use in the abstract, so a purple-pipe agreement is worth negotiating into the siting deal rather than after it. An on-site treatment plant is a different size of decision, since it comes with operators, permits, and discharge limits, and it pays only where campus volumes justify a real facility. Report water use effectiveness next to PUE, because a site that improved its PUE by moving to evaporative cooling moved a cost rather than removing one.

Videos
The Big Data Center Water ProblemAsianometry · 100k+ views
Data Center Water is a DistractionKyle Hill · 100k+ views
How data centers stay cool while reducing water usage demandsWATE 6 On Your Side · under 1k views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · Cooling Water Efficiency Opportunities for Federal Data Centers (US Department of Energy Federal Energy Management Program)

A data center converts almost all the electricity it draws into low-grade heat, and normally throws it away. Heat reuse captures it instead and sells it, usually into a district heating network. Air-cooled halls produce 30–40 °C return air, which is too cool to use directly and needs a heat pump to lift it to the 60–80 °C a network wants. Liquid cooling changes this materially: a direct-to-chip return at 45–50 °C needs far less lifting, and some networks now accept it with a small heat pump or none at all.

Strengths & weaknesses

The heat exists whether or not anyone uses it, so the marginal carbon benefit of displacing a gas boiler is real and the revenue, while modest, is genuine. In cities with existing district heating this is a straightforward commercial arrangement. The obstacles are geography and mismatch. A network has to exist within a few kilometers, the operator has to want a supply whose availability depends on someone else's business, and heat demand is seasonal while the data center runs flat. Contracts also constrain the data center's own operation, since it must now keep return temperature within the network's specification.

When to use

Pursue heat reuse where a district network exists nearby and the local regulator or planning authority values it, which describes most of northern Europe. It fits liquid-cooled facilities much better than air-cooled ones, so it belongs in the design conversation when a hall is converting anyway. Do not build a business case around the heat revenue; it is small next to compute revenue. Treat it as a planning and community asset, since in several jurisdictions offering heat has become part of getting permission to build at all.

Key numbers

Nearly all electricity drawn becomes low-grade heat · air-cooled return air at 30–40 °C needs a heat pump to reach the 60–80 °C district networks want · liquid-cooled return at 45–50 °C needs far less lift · viable only within a few kilometers of an existing network · heat demand is seasonal while data center output is constant.

Examples

Google's heat recovery project in Hamina, Finland, feeding a local network; Fortum's Espoo scheme taking heat from Microsoft datacenters; Stockholm Data Parks, which built the model; Danish and Dutch planning rules that now expect heat reuse from new facilities.

Economic profile

Two capital items decide this: the heat pump that lifts return temperature to what the network wants, and the pipe run to the connection point. Pipe cost scales with distance, which is why the practical limit is a few kilometers and why a site outside that radius does not improve with a bigger heat pump. Lift is the recurring cost, so air-cooled return at 30–40 °C against a 60–80 °C network target is the expensive case and liquid-cooled return at 45–50 °C is the cheap one. Heat revenue is small next to compute revenue either way, so it should not carry the business case. The value that does appear is in permitting: Danish and Dutch rules now expect heat reuse from new facilities, so the connection is part of the price of building at all. If a hall is being converted to liquid cooling anyway, price the connection then, since adding it later means opening the plant a second time.

Videos
Finland’s Big Idea: Turning Data Center Waste Into HeatBloomberg Television · 100k+ views
Google’s first-ever heat recovery project for neighbourhoods in FinlandGoogle · 10k+ views
Waste heat from data centresFortum · 1k+ views
Further reading

AI: Five charts that put data-centre energy use - and emissions - into context (Carbon Brief) · Data center waste heat for district heating networks: A review (Renewable and Sustainable Energy Reviews, via Aalto University)

Class VI

Facility & siting

building type, ownership model, and location5 systems

A hyperscale campus is a purpose-built site of several buildings, each 30–150 MW, sharing a substation, a water supply, and a security perimeter. Design is standardized and repeated: the same hall, the same electrical block, the same mechanical arrangement, built again and again so construction becomes a manufacturing exercise rather than a bespoke project. The operator owns the compute, so the building can be optimized against its own hardware rather than against a generic tenant specification.

Strengths & weaknesses

Repetition drives everything good about it: cost per megawatt falls, construction schedules compress to 12–24 months per building, and operating practice transfers between sites. Because the operator controls the IT, it can run hot aisle containment at aggressive set points, standardize on one cooling architecture, and design power around its own rack specification. The downsides are concentration and inflexibility. A campus is a very large bet on one location's power, water, and politics, and a standardized design that assumed 30 kW racks is expensive to convert when the next generation needs 130 kW.

When to use

This model belongs to operators with enough demand to fill several buildings and enough control over the hardware to design for it. If you are that buyer, the leverage comes from standardizing early and building repeatedly, and from securing power and water before land. If you are not, colocation or a build-to-suit lease gets similar economics without the balance sheet. The genuine risk to plan for is generational: assume the density specification will change during the campus's life, and pipe for liquid cooling even where the first buildings do not need it.

Key numbers

Individual buildings typically 30–150 MW, campuses 100 MW to over 1 GW · construction of 12–24 months per building once the design repeats · standardized design cuts cost per MW substantially against bespoke construction · US data center electricity reached about 176 TWh in 2023 and is projected to roughly double or triple by 2028 · retrofitting a hall designed for air to liquid is a major cost.

Examples

Meta's Prineville and Odense campuses; Microsoft's Boydton and San Antonio sites; the Amazon campus in Indiana built for Anthropic workloads; Google's Council Bluffs campus, one of the largest single sites in the US.

Economic profile

Hyperscale economics come from repetition and from owning the whole stack, so a saved dollar per watt shows up dozens of times. The dominant risk has shifted from construction to inputs: power availability, transformer lead time, and now GPU supply set the schedule, not concrete. That is why operators are signing power purchase agreements and reserving equipment years ahead, and why "powered land" trades at a premium over ordinary industrial land.

Videos
Microsoft reveals its MASSIVE data center (Full Tour)CNET Highlights · 500k+ views
No Nvidia Chips Needed! Amazon’s New AI Data Center For Anthropic Is Truly MassiveCNBC · 1m+ views
The Full Tour: Wisconsin’s First Hyperscale Data Center ConstructionRCNFRD · 1k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · Key Questions on Energy and AI (IEA)

Prefabricated modular means building the data center in a factory and shipping it. Scope varies: skid-mounted power or cooling plant that arrives tested and only needs connecting, all-in-one containerized halls with racks already installed, or full modular buildings assembled from repeated volumetric units. The common thread is moving work from a congested site with a scarce skilled trade to a controlled factory where the same assembly is built repeatedly.

Strengths & weaknesses

Schedule is the main product. Factory work runs in parallel with site work rather than after it, and deployment times fall from 18–24 months to 6–12. Quality is more consistent because commissioning happens on a production line, and the approach suits places where skilled electrical labor is unavailable. Against that, cost per megawatt is usually higher than a well-executed stick-built project, module dimensions are constrained by what a truck can carry, and the design is fixed at order time, so late changes are expensive. Some designs lock the buyer into one vendor's mechanical and electrical ecosystem.

When to use

Choose prefabricated where speed matters more than capital cost, where site labor is scarce, at remote or edge locations, and for capacity that must be added in increments rather than all at once. Prefabricated power and cooling skids inside a conventional building are the most widely useful version, capturing most of the schedule benefit without committing the whole building to a module vendor. Stick-build where the site has good labor, the program is large enough to amortize a bespoke design, and cost per megawatt is the metric being optimized.

Key numbers

Deployment in roughly 6–12 months against 18–24 for conventional construction · module size limited by road transport, typically 12–14 m long · factory commissioning reduces on-site work and rework · cost per MW usually above a well-executed stick-built project · design locked at order, so late changes are expensive.

Examples

Vertiv and Schneider Electric prefabricated power and cooling modules; Aligned's modular builds in Ohio; the containerized data centers that Microsoft and Google used a decade ago and that returned in a different form for AI capacity; edge modules deployed at cell tower sites.

Economic profile

A buyer pays more per megawatt than a well-executed stick-built project and gets capacity in 6–12 months instead of 18–24. Whether that premium is worth it comes down to what a year of earlier revenue is worth: where tenants are waiting on power it usually clears the premium comfortably, and where the capacity would sit idle it does not. Local labor moves the comparison as much as the vendor's price does, because where electrical trades are scarce the conventional schedule runs at the long end of 18–24 months or past it, and the saving is larger than the headline numbers suggest. The cheapest way to capture most of it is prefabricated power and cooling skids inside a conventional building, which keeps the schedule benefit without committing the whole building to one vendor's mechanical and electrical ecosystem. The item to price carefully is the design freeze at order, since changes after it are made at factory change-order rates and module dimensions are fixed at 12–14 m by road transport. If the program is large enough to amortize a bespoke design and cost per MW is the metric being optimized, stick-build.

Videos
Prefabricated Modular Data Center Tour | Vertiv™ SmartMod™ MaxVertiv · 5k+ views
The Future of Prefabricated Modular Data Centers | Schneider ElectricSchneider Electric · 10k+ views
A smarter, faster way to build AI-ready data centers | Vertiv™ OneCoreVertiv · 10k+ views
Further reading

Exploring the Efficiency of Renewable Energy-based Modular Data Centers at Scale (arXiv) · Types of Prefabricated Modular Data Centers, White Paper 165 (Schneider Electric)

Colocation is renting space, power, and cooling in someone else's building while owning the servers yourself. Retail colocation sells by the cabinet or the cage, with the operator providing everything up to the rack. Wholesale sells whole halls or buildings, typically 1 MW and up, with the tenant taking on more of the fit-out and operations. Contracts are priced primarily on committed kilowatts rather than on floor area, which tells you what the operator is really selling.

Strengths & weaknesses

It converts capital into operating expense, gives access to carrier-dense interconnection points that would be impossible to replicate, and removes the need to run a critical facility. Deployment takes weeks instead of years. The costs are per-kilowatt price and constraint. Colocation power is more expensive than self-build at scale, halls built for 5–10 kW racks often cannot take modern density, and liquid cooling in a shared hall requires the operator's cooperation on piping, water treatment, and leak response. Long contracts also lock in a density specification that may age badly.

When to use

Colocate when the requirement is under roughly 5–10 MW, when interconnection to many networks matters, or when speed matters more than unit cost. It is also the right way to enter a new geography before committing to a build. Check three things before signing: the hall's actual per-rack power and cooling limit, whether liquid cooling is supported and on what terms, and how power is billed, since metered against committed changes the economics substantially. Above about 10 MW with a stable forecast, build or lease a whole facility instead.

Key numbers

Retail sold by cabinet, wholesale from about 1 MW · priced on committed kW rather than floor area · deployment in weeks against years for a build · legacy halls commonly cap at 5–15 kW per rack, which excludes current AI hardware · liquid cooling support varies widely and is a contract term, not a given.

Examples

Equinix and Digital Realty in the retail and interconnection market; CyrusOne, Vantage, and QTS in wholesale; the Ashburn and Slough interconnection clusters, where colocation exists mainly for who else is in the building; colocation halls now retrofitting rear-door heat exchangers to take AI tenants.

Economic profile

Colocation is sold by the committed kilowatt, so the contract is a power lease with a building attached and floor area barely enters the price. That is why the crossover with self-build sits around 5–10 MW: below it, paying the operator's margin costs less than building and staffing a critical facility, and above it self-build wins on unit cost as long as the demand forecast holds. Two contract terms move the economics more than the headline rate does. Metered against committed billing decides who carries the cost of unused capacity, and a hall capped at 5–15 kW per rack can quote an attractive price per kW that means nothing if the hardware you want does not fit in it. Liquid cooling is also a contract term rather than a given, so a tenant who expects to run AI hardware in year three should get piping, water treatment, and leak response written in at signing. The one thing colocation sells that self-build cannot is interconnection: in Ashburn or Slough part of the rent is for who else is in the building, and no capital budget substitutes for that.

Videos
What is Colocation & How Does It Work?Interxion: A Digital Realty Company · 500k+ views
The 4 Types of Data Centers Explained | Enterprise, Colocation, Hyperscale & EdgeData Center Resources · 5k+ views
Understanding Colocation Data CentersProvision Networks · 1k+ views
Further reading

Uptime Institute Global Data Center Survey 2025 (Uptime Institute) · Colocation Data Centers (Better Buildings, US Department of Energy)

An edge data center trades scale for proximity. Instead of one large facility a thousand kilometers away, compute sits in a small enclosure near the users: a cabinet in a cell tower compound, a container in a retail parking lot, a room in a regional exchange. Capacity is typically 50 kW to a few megawatts. The purpose is latency, local data handling, and bandwidth cost, since processing video near where it is generated is far cheaper than backhauling it.

Strengths & weaknesses

Latency of a few milliseconds instead of tens is the product, and for real-time control, autonomous systems, and interactive video that difference is the whole application. Local processing also keeps regulated data inside a jurisdiction. The costs are efficiency and operations. Small sites have poor PUE, typically 1.5–2.0, because the fixed overhead of cooling and power conversion does not shrink with load. Nobody is on site, so everything needs remote hands and remote power cycling, physical security is weaker, and managing hundreds of small sites costs more per kilowatt than managing one large one.

When to use

Deploy at the edge when latency or data locality genuinely requires it, and be strict about that test, because most workloads do not. Content caching, industrial control, retail analytics, and telecom network functions are the durable cases. Design for zero site visits: switched PDUs, out-of-band management, and sealed cooling. Where the requirement is only capacity rather than proximity, a regional colocation facility is cheaper, more efficient, and far easier to run.

Key numbers

Typical capacity 50 kW to a few MW · latency in the low single-digit milliseconds against tens for a distant region · PUE commonly 1.5–2.0 because overhead does not scale down · no staff on site, so remote management is mandatory · per-kilowatt operating cost is higher than any other facility type here.

Examples

Cell tower compounds hosting telecom network functions; EdgeConneX and DataBank regional sites; content delivery caches inside internet exchanges; industrial edge cabinets on factory floors running control and vision workloads.

Economic profile

Almost every cost here is per site, and the sites are small, so overhead that a large facility spreads over tens of megawatts gets spread over 50 kW. It shows up twice: PUE of 1.5–2.0 against roughly 1.1 at a large hyperscale site, and a per-kilowatt operating cost higher than any other facility type on this sheet. The case therefore has to come from somewhere other than unit cost, and there are only two places it comes from. One is bandwidth, since processing video where it is generated instead of backhauling it is a straightforward comparison of transport cost against the per-kilowatt premium. The other is revenue that does not exist without low latency, which is a much harder claim and the one most edge business cases lean on without testing. If the requirement is capacity rather than proximity, a regional colocation hall is cheaper on every line, and designing for zero site visits is the main lever an operator has on what remains.

Videos
What are edge data centres and why are they essential for 5G?Business Standard · 10k+ views
Micro Data Centers for Edge ComputingTripp Lite · 1k+ views
Edge data centersCAREL · 5k+ views
Further reading

Energy Efficient Deployment and Orchestration of Computing Resources at the Network Edge: a Survey on Algorithms, Trends and Open Challenges (arXiv) · A Comprehensive Survey of Micro Datacenter: Current Technologies and Future Possibilities (Frontiers of Computer Science, via Shanghai Jiao Tong University)

Site selection used to weigh land, fiber, tax, and climate. It now starts with one question: how many megawatts can this location deliver and when. Everything else is secondary, because a site with cheap land, good fiber, and no interconnection date is not a site. The rest of the screen covers water availability and permitting, climate for economization, natural hazard exposure, latency to the target users, construction labor, and the local political appetite for a very large industrial load.

Strengths & weaknesses

Getting this right is the highest-leverage decision in the whole project: it fixes power price for decades, decides whether the facility can economize, and sets how much water the design can use. A well-chosen site makes an ordinary design perform well. The difficulty is that the good sites are known and contested. Powered land with an executed interconnection agreement trades at a large premium, several jurisdictions have introduced moratoria or new rules on large loads, and the utility's answer often depends on a transmission upgrade whose schedule nobody controls.

When to use

Screen on power first, water second, and everything else third. Ask the utility for a real capacity date rather than an expression of interest, and evaluate whether a flexible-load or curtailable agreement would move that date. Check climate against the cooling architecture: a site that supports dry cooling year round removes the water question entirely. Talk to the community before the announcement rather than after, since local opposition has become a genuine schedule risk. And be honest about hazards, because insurance and downtime cost more than the land ever will.

Key numbers

Interconnection studies of one to four years, and longer where a transmission upgrade is needed · powered land with an executed agreement carries a substantial premium over raw industrial land · climate determines economizer hours and therefore both PUE and water use · several US states and European jurisdictions have added rules or moratoria on large data center loads · latency to major population centers sets which workloads the site can serve.

Examples

Northern Virginia, where transmission constraint rather than land now limits growth; ERCOT's large-load queue in Texas; Ireland's effective moratorium in the Dublin region; the shift toward Ohio, Georgia, and the Midwest, driven mainly by available substation capacity.

Economic profile

Site selection fixes a cost structure for decades, which makes the premium on powered land rational rather than speculative: an executed interconnection agreement removes one to four years of queue, and the developer paying that premium is buying a date rather than acreage. Power is 15–30% of total cost of ownership, so a cheaper rate at a site that connects three years later is usually the worse deal, and the arithmetic only reverses for a buyer with no urgency. Climate is the other durable line item, because economizer hours set both compressor energy and water use for the life of the facility, and no later design change recovers a site that has neither. Two items are worth pricing that projects often treat as free: a contribution in aid of construction toward the utility's network upgrade, which lands as capital, and a curtailable or flexible-load agreement, which can move the connection date forward and should be modeled against the revenue given up during curtailment. Local opposition belongs in the schedule risk, since Ireland's Dublin-region restrictions and new rules in several US states arrived after developers had already bought land.

Videos
Telx - Ideal Data Center Site Selection | Schneider ElectricSchneider Electric · under 1k views
Further reading

Key Questions on Energy and AI (IEA) · Data Center Energy Infrastructure: Federal Permit Requirements (Congressional Research Service, via EveryCRSReport)

Class VII

Racks & operations

standards, interconnect, and how it is run5 systems

The 19-inch rack dates from railway signaling equipment and has survived largely by inertia. The Open Compute Project's Open Rack replaced it for hyperscale use: a 21-inch equipment opening in the same 600 mm floor footprint, a shared DC busbar down the back so servers have no individual power supplies or cords, and centralized power shelves and fans. Open Rack v3 raised the busbar to 48 V, added support for far higher rack power, and standardized the mechanical interfaces for liquid cooling manifolds.

Strengths & weaknesses

Centralizing power conversion is more efficient than dozens of small supplies, removes a large number of cables and connectors, and makes rack-level redundancy simpler. The wider opening gives more room for heatsinks and airflow, and blind-mate busbar connections make servers faster to install and remove. The cost is ecosystem: Open Rack equipment is bought from a smaller set of suppliers, it does not fit a standard 19-inch cabinet, and the design is aimed at operators who buy hundreds of racks of one configuration. For a mixed enterprise fleet it is the wrong shape entirely.

When to use

Adopt Open Rack when buying at hyperscale volume with a uniform hardware fleet, since the efficiency and serviceability gains multiply. It is also increasingly the practical choice for dense AI deployments, because current accelerator racks are designed around this form factor and the liquid cooling manifolds assume it. Stay with 19-inch racks for enterprise and colocation, where equipment comes from many vendors and cabinets have to accept whatever arrives. Whatever the choice, confirm floor loading early; a fully populated liquid-cooled rack can exceed 1,500 kg.

Key numbers

21-inch equipment opening within the same 600 mm floor pitch as a standard rack · 48 V DC busbar in Open Rack v3, replacing per-server power supplies · centralized power shelves are more efficient than distributed supplies · a fully populated liquid-cooled rack can exceed 1,500 kg, well beyond typical raised-floor ratings · specifications are published openly rather than licensed.

Examples

Meta, which created the Open Compute Project and runs Open Rack across its fleet; NVIDIA's GB200 NVL72, built on an Open Rack-derived form factor; Microsoft's contributions of its own rack and cooling designs to OCP; the OCP specification library, which is the reference for interfaces and power.

Economic profile

The savings are per rack and small, so they only count multiplied out: one fewer power supply in every server, one fewer conversion stage, and blind-mate busbars that cut install and swap labor. Across hundreds of racks of one configuration that adds up; across twenty mixed cabinets it does not cover the cost of running a second supply chain. The specifications are published rather than licensed, so no royalty is involved, but the supplier set is smaller and the equipment does not fit a standard 19-inch cabinet, which makes the switch close to all-or-nothing for a hall. For AI buyers the decision has largely been made upstream, since current accelerator racks such as the GB200 NVL72 ship in an Open Rack-derived form factor and the liquid cooling manifolds assume it. The cost that surprises people is structural: a fully populated liquid-cooled rack can pass 1,500 kg, and reinforcing a floor to take that is a building expense that usually costs more than the rack standard saves. If the fleet is mixed and arrives from many vendors, stay on 19-inch and spend the money elsewhere.

Videos
OCP Gear Explained - What's inside an OCP Rack?Open Compute Project · 5k+ views
Teaser of the new OCP Academy course series on Open Rack ORv3Open Compute Project · under 1k views
Intro to Rack & Power TrackOpen Compute Project · under 1k views
Further reading

Toward Next-Generation AI Data Centers: Power Delivery Architecture Shifts, Emerging Technologies, and Challenges (arXiv) · Facebook announces next-generation Open Rack frame (Engineering at Meta)

Inside a data center, almost everything past a few meters is optical. Structured cabling runs single-mode or multimode fiber from rack to row to spine in a fixed hierarchy, and pluggable transceivers at each end convert electrical signals to light and back. Rates have climbed from 10G through 100G and 400G to 800G per port, and AI clusters have made the network a first-order design problem rather than a utility, because training performance depends on how fast thousands of accelerators can exchange gradients.

Strengths & weaknesses

Fiber carries far more bandwidth per strand than copper, over distances copper cannot reach, with no crosstalk and a very long service life; the cabling plant usually outlives several generations of transceiver. Structured design means adding capacity is a patch rather than a pull. The costs are optics and power. Transceivers are a large share of network capital cost, they consume real power at high rates, and at 800G and above the pluggable module's power becomes a rack-level concern, which is what is pushing co-packaged optics. Cleanliness matters more than people expect; a contaminated connector is the most common fault in the plant.

When to use

Design a structured cabling plant once, generously, and treat it as a 15-year asset while transceivers turn over every few years. Use single-mode where reach or future rate is uncertain, since it costs a little more to install and removes the distance ceiling. In AI clusters, design the network topology alongside the compute rather than after it, and check the optics power budget per rack, because at 800G it is no longer negligible. For short intra-rack links, direct-attach copper remains cheaper and lower power.

Key numbers

Port rates now 400G and 800G, with 1.6T in development · single-mode reaches kilometers, multimode tens to hundreds of meters · transceivers are a large share of network capital and consume meaningful rack power at high rates · cabling plant typically lasts 15 years across several transceiver generations · connector contamination is the most common cause of link faults.

Examples

Spine-and-leaf fabrics in every hyperscale facility; InfiniBand and Ethernet fabrics inside AI training clusters, where interconnect bandwidth limits scaling; co-packaged optics programs from Broadcom and NVIDIA aimed at the transceiver power problem; TIA-942 structured cabling practice, which defines the hierarchy most facilities follow.

Economic profile

The money splits across two very different lifetimes. The cabling plant is a 15-year asset whose cost is mostly labor, and pulling fiber into a live hall later costs far more than over-specifying it at build, which is why single-mode is usually the right call even where multimode would reach today. Transceivers are the other half: they turn over every few years, they are a large share of network capital, and at 800G their power draw becomes a rack-level line item. That matters more here than in most buildings, because when power is the binding constraint every watt spent in an optical module is a watt not sold as compute, which is the whole argument behind co-packaged optics. Two smaller items are worth budgeting honestly. Direct-attach copper is cheaper and lower power for short intra-rack links, and connector cleaning discipline is one of the cheapest reliability measures in the building, since contamination causes more link faults than anything else.

Videos
Structured Cabling for Large Data Centers: An Inside Look (Ep. 49)CABLExpress · 10k+ views
Fiber Optic Cabling Solutions for Data Centers | FSFS_com · 1k+ views
What is structured cabling in networking? (Structured Data Cabling)NM Cabling Solutions · 50k+ views
Further reading

2024 United States Data Center Energy Usage Report (Berkeley Lab) · Co-packaged optics (CPO): status, challenges, and solutions (Frontiers of Optoelectronics)

A building management system runs the mechanical and electrical plant: chillers, pumps, air handlers, generators, and switchgear, with alarms and set points. Data center infrastructure management sits alongside it and tracks the IT side: what is in each rack, what each circuit draws, what the inlet temperatures are, and how much power and cooling capacity remains where. The two together answer the question every operator has to answer daily, which is whether the next deployment fits.

Strengths & weaknesses

Done well, DCIM turns capacity planning from an argument into a calculation, catches stranded capacity, and prevents the common failure of a hall that is out of power in one row and empty in another. Asset tracking cuts the time to find and service equipment. The weakness is data quality. A DCIM whose asset records drift from reality is worse than none, because people trust it, and keeping records accurate needs process discipline that many organizations do not sustain. Deployments are also frequently oversold: the software is easy to buy and the operational change is what actually delivers the value.

When to use

Deploy DCIM once the facility is big enough that nobody can hold it in their head, roughly above a few hundred racks or wherever multiple teams share capacity. Start with the measurements that drive decisions, which are per-circuit power and per-rack inlet temperature, and expand from there rather than trying to model everything on day one. Insist on automated data collection wherever possible, since manual entry is where accuracy dies. Keep the building management system on a separate, tightly controlled network, because it can operate plant.

Key numbers

Per-circuit power and per-rack inlet temperature are the two measurements that drive most decisions · stranded capacity of 10–30% is common in facilities without good instrumentation · asset record accuracy decays quickly without automated collection · building management systems control plant and therefore sit inside the security perimeter · payback comes from deferred capacity rather than from energy alone.

Examples

Nlyte, Sunbird, and Schneider EcoStruxure IT in enterprise and colocation; hyperscale operators running their own internal tooling rather than commercial DCIM; colocation providers exposing per-cabinet power data to tenants as a product feature.

Economic profile

The payback is deferred capital rather than energy. A facility with 10–30% of its power or cooling stranded in the wrong row has already paid for capacity it cannot sell, and recovering part of that is worth far more than the software costs, because the alternative is building megawatts that already exist. That comparison is why DCIM keeps getting bought and why it so often disappoints: the license is the cheap part, and the expensive part is the process discipline that keeps asset records true, which never appears on the quote. A good rule of thumb is to budget more for automated data collection than for the software itself, since manual entry is where accuracy dies and an inaccurate system is worse than none because people act on it. Start with per-circuit power and per-rack inlet temperature, which is where nearly all the capacity decisions come from, and expand only once those are trustworthy.

Videos
Data Center Infrastructure Management (DCIM) ExplainedAnixter · 50k+ views
What is DCIM? - Data Center Infrastructure Management ExplainedFLUIXAI · 10k+ views
Why DCIM Software is a Game ChangerData Center News · 1k+ views
Further reading

Intelligent Monitoring of Data Center Physical Infrastructure (Applied Sciences) · Avoiding Common Pitfalls of Evaluating and Implementing DCIM Solutions, White Paper 170 (Schneider Electric)

Uptime Institute's Tier classification describes how much of a facility can fail or be maintained without stopping the IT load. Tier I is a single path with no redundancy. Tier II adds redundant components. Tier III is concurrently maintainable: any element can be taken out of service for work with the load still running. Tier IV is fault tolerant: an unplanned failure of any single element does not affect the load. Commissioning is the separate discipline of proving the design actually behaves that way, in five levels from factory testing through integrated systems testing under simulated failure.

Strengths & weaknesses

The Tier system gives buyers and designers a shared vocabulary, and a certified design is a genuine commercial signal rather than a marketing claim. Level 5 integrated systems testing, where the whole facility is run at load and faults are deliberately introduced, is the single most effective way to find design and construction errors before a tenant does. The costs are capital and misuse. Tier IV can cost 30–50% more than Tier III for redundancy most workloads no longer need, and the label is widely applied loosely, with "Tier III design" claimed where nothing was ever certified.

When to use

Choose the tier from how the application handles failure. Distributed cloud and AI training workloads that tolerate node and even site loss do not need Tier IV; several hyperscale designs are deliberately below Tier III at the facility level because resilience lives in the software. Enterprise systems with no failover still need concurrent maintainability, and Tier III is the usual answer. Whatever the tier, commission properly and insist on integrated systems testing under load, because an untested redundant design is a redundant design on paper only.

Key numbers

Tier III is concurrently maintainable; Tier IV is fault tolerant against any single unplanned failure · Tier IV typically costs 30–50% more than Tier III · commissioning runs in five levels, ending with integrated systems testing under load · most human-error outages trace to procedures rather than equipment · certification applies to the design, the constructed facility, or the operations, and the three are separate.

Examples

Uptime Institute Tier certifications held by colocation providers as a sales credential; hyperscale designs that deliberately reduce facility redundancy because the application fails over between sites; integrated systems testing catching control-sequence errors that no component test would have found.

Economic profile

Redundancy multiplies the most expensive part of the building, so the tier choice moves total capital more than almost anything else here: Tier IV typically costs 30–50% more than Tier III for the same IT capacity. Whether that is money well spent depends on where resilience already lives, and for a workload that fails over between sites in software the operator is paying for it twice. That is why several hyperscale designs sit deliberately below Tier III at the facility level, and why an enterprise application with no failover still buys Tier III concurrent maintainability. A colocation tenant should ask what a tier claim actually covers, since certification of the design, the constructed facility, and the operations are three separate things, and "Tier III design" gets claimed where nothing was certified. Commissioning is the opposite kind of spend: it lands at the end when the schedule is already late, it costs far less than the redundancy it is checking, and level 5 integrated systems testing under load is the cheapest place in the program to find a design or construction error. If it gets cut, a tenant finds the error instead.

Videos
Uptime Data Center Tier Levels - The Gold StandardData Center News · 1k+ views
The 5 Levels of Data Center Commissioning (Explained)Five Nines · 5k+ views
CertMike Explains Data Center TiersMike Chapple · 5k+ views
Further reading

Uptime Institute Global Data Center Survey 2025 (Uptime Institute) · Uptime Institute Tier Classification System (Uptime Institute)

Power usage effectiveness is total facility energy divided by IT energy. A PUE of 1.5 means half a watt of overhead for every watt of computing. Water usage effectiveness is the same idea in liters per kilowatt-hour of IT load. Both are ratios, which is their strength and their weakness: they compare a facility against itself over time honestly, and they compare two facilities against each other only if measured the same way, at the same boundary, over the same period.

Strengths & weaknesses

PUE drove a genuine decade of improvement, because it gave operators one number to manage and a clear target. It is easy to measure once the metering exists, and annualized PUE is hard to game. The weaknesses are what it excludes. PUE says nothing about whether the IT is doing useful work, so replacing servers with more efficient ones makes PUE worse while cutting total energy. It ignores water entirely, which is why WUE exists and why a site can improve PUE by moving to evaporative cooling while making its water position worse. Industry average PUE has been flat near 1.5 for six years.

When to use

Measure annualized PUE and WUE at consistent boundaries and use them to track your own facility. Do not use them to rank facilities in different climates, and do not let a PUE target drive a decision that raises total resource use. For anything about compute efficiency, use a work-per-energy measure instead, since PUE deliberately says nothing about it. When comparing vendor claims, ask what was included, over what period, and at what load, because a design-day figure at full load is a different number from an annualized one.

Key numbers

PUE is total facility energy divided by IT energy; WUE is liters of water per kWh of IT energy · industry weighted average PUE was about 1.54 in 2025, roughly unchanged for six years · large hyperscale sites report trailing PUE around 1.1 · a partly loaded facility has a worse PUE than the same facility at full load · annualized measurement is the only fair basis for comparison.

Examples

The Green Grid's original PUE definition and its later standardization in ISO/IEC 30134; Google's published fleet-wide trailing PUE, among the lowest reported; Uptime Institute survey data showing the industry average stalled; European regulations now requiring reporting of both energy and water for large facilities.

Economic profile

Overhead shows up twice, as an electricity bill and as capacity that cannot be sold as compute. At 10 MW of IT load, moving from a PUE of 1.5 to 1.4 removes 1 MW of overhead (15 MW total against 14), and where power is the binding constraint the freed megawatt is usually worth more than the energy saved. That is what makes containment the best-returning spend in a legacy room, since it recovers 20–40% of cooling energy for very little capital. Past that the curve flattens, which is why the weighted industry average has sat near 1.54 for six years while new hyperscale sites report around 1.1: most of the fleet is old buildings where the constraint is the building itself. Anyone putting these ratios into a financial model should watch two things. A PUE target can be met by moving to evaporative cooling while the water bill and the local water politics get worse, and PUE improves when the servers get less efficient, so the ratio should never stand in for compute efficiency or be used to rank sites in different climates.

Videos
What is PUE? - Data Center EfficiencyFLUIXAI · 5k+ views
What is PUE, Why Important ? Power Usage Effectiveness Explained | Data Center Efficiency SimplifiedCyber Project Manager EN · under 1k views
What is PUE Finallknoxie7 · 1k+ views
Further reading

Power usage effectiveness (Google Data Centers) · Electrical Efficiency Measurement for Data Centers, White Paper 154 (Schneider Electric)

Glossary

Terms that show up in the system explorer and are not obvious from outside the field. Numbers are typical values, not specifications.

TermWhat it means
Adiabatic assistSpraying a fine mist onto a dry cooler's coil so that evaporation boosts its capacity on hot days. It lets the plant be sized for average conditions instead of the design day, and it uses water only for the handful of hours that need it.
Aisle containmentA physical barrier that keeps cold supply air and hot exhaust air from mixing, by enclosing either the cold aisle or the hot one. It typically cuts cooling energy 20–40% in a room that had none, and it is the cheapest large efficiency gain in air cooling.
BlowdownWater deliberately drained from a cooling tower to stop dissolved solids concentrating as evaporation removes pure water. It adds to total water draw and its quality is regulated, so it cannot simply be discharged.
BuswayAn enclosed conductor bar run above the racks, into which tap-off boxes plug anywhere along its length. It turns adding a circuit into a plug-in operation, and its rating is chosen at design time and expensive to change.
CDUA coolant distribution unit: the heat exchanger, pumps, and controls that separate the building's water from the clean treated fluid circulating through cold plates or rear doors. It lets the two loops run at different temperatures, pressures, and chemistries.
Cold plateA metal block with internal channels, clamped directly onto a processor package, through which coolant flows. Because the thermal path is short, the coolant can be much warmer than air would need to be, which is what makes chiller-free operation possible.
CommissioningProving that what was built behaves the way the design intended, in levels from factory testing to integrated systems testing where the whole facility runs at load and faults are deliberately introduced. It is where control-sequence errors get found, and skipping it is how a redundant design turns out not to be.
CRAC and CRAHA computer room air conditioner has its own refrigeration circuit; a computer room air handler is a coil and fan fed with chilled water from a central plant. The distinction decides where the compressor lives and therefore how efficiently the whole site can run.
Dielectric fluidA liquid that does not conduct electricity, so electronics can be immersed in it or it can flow through sealed loops touching live parts. The good ones for two-phase work are fluorocarbons, which is the reason PFAS regulation matters to cooling design.
Double conversionA UPS topology that rectifies incoming AC to DC and inverts it back, so the load is always fed from the inverter and never sees the utility waveform. Transfer time on a utility failure is zero because there is no transfer.
Dry bulb and wet bulbOrdinary air temperature, and the lowest temperature reachable by evaporating water into that air. Evaporative equipment approaches the wet bulb and dry equipment approaches the dry bulb, which is why humid climates suit dry rejection and arid ones suit evaporative.
EconomizationUsing outdoor conditions to do the cooling instead of running a compressor, either by bringing in outside air or by bypassing the chiller when the tower can make cold enough water. Compressors are the largest mechanical load, so free hours translate directly into efficiency.
Immersion coolingSubmerging whole servers in a bath of dielectric fluid so every component is cooled and no fans are needed. Single-phase circulates the fluid past the boards; two-phase lets it boil on the hot parts and condense on a coil above.
HyperscaleAn operator large enough to build standardized facilities repeatedly and to own the compute inside them, so the building can be designed around its own hardware. Individual buildings run 30–150 MW and campuses can exceed a gigawatt.
N+1 and 2NRedundancy notation. N is what the load needs; N+1 adds one spare unit; 2N duplicates the whole system on independent paths. N+1 usually gives concurrent maintainability, and 2N is what survives a failure during maintenance.
Open RackThe Open Compute Project's rack standard: a 21-inch equipment opening in a standard floor footprint, with a shared DC busbar replacing individual server power supplies and cords. It suits operators buying hundreds of identical racks and does not fit a mixed enterprise fleet.
PDUA power distribution unit. In the room it is the transformer and panel that feeds the racks; in the rack it is the metered strip the servers plug into. Per-outlet metering on the rack version is what capacity planning actually runs on.
PFASPer- and polyfluoroalkyl substances, the fluorinated chemistry behind most two-phase cooling fluids. They persist in the environment, 3M announced an exit from PFAS manufacture by the end of 2025, and EU restrictions are advancing, which puts two-phase supply chains in question.
PUEPower usage effectiveness: total facility energy divided by IT energy. A PUE of 1.5 means half a watt of overhead per watt of computing. It compares a facility against itself honestly and compares different facilities only if the boundary and period match.
Quick disconnectThe dripless coupling that lets a liquid-cooled server be removed without draining the loop. It is what makes cold-plate cooling serviceable, and it is also the component that turns a routine swap into a wet operation.
Raised floorA structural floor on pedestals with a plenum beneath, historically used to deliver cold air through perforated tiles and to route cabling. It caps air delivery at roughly 5–15 kW per rack and rarely carries the weight of a populated liquid-cooled rack.
Rear-door heat exchangerA water-cooled coil that replaces a rack's back door, so exhaust air is cooled before it leaves the cabinet. From the room's point of view the rack produces no heat, which adds density without touching the hall's air handling.
Stranded capacityPower or cooling that exists but cannot be used, because it is in the wrong row, behind the wrong breaker, or blocked by a busway rating. It commonly runs 10–30% in facilities without good instrumentation, and finding it is cheaper than building more.
Tier classificationUptime Institute's scheme for how much of a facility can fail or be maintained without stopping the load. Tier III is concurrently maintainable, Tier IV is fault tolerant against any single unplanned failure, and the difference is 30–50% of capital.
Two-phase and single-phaseWhether the coolant boils. Single-phase stays liquid and carries heat by rising in temperature; two-phase evaporates, which absorbs far more heat per unit flow and holds the surface near the boiling point. Two-phase is thermally better and commercially constrained by fluid chemistry.
WUEWater usage effectiveness: liters of water consumed per kilowatt-hour of IT energy. It exists because a site can improve its PUE by switching to evaporative cooling while making its water position considerably worse.

How to choose data center infrastructure

Two questions decide almost everything else: how many kilowatts per rack, and where does the heat go. Density picks the cooling architecture, cooling architecture picks the loop temperature, and loop temperature decides whether the site can reject heat dry or has to evaporate water. Get those three in order and most of the rest follows. Get them out of order and you build a hall that cannot take the hardware you bought.

Density is the fork in the road

Air can carry a bounded amount of heat out of a cabinet. In practice a well-contained air-cooled hall tops out somewhere between 20 and 40 kW per rack, and beyond that you are moving so much air that fan power and acoustics become the problem. AI training racks are already at 80–140 kW and the roadmaps go higher. That is not an incremental change to an air-cooled design; it is a different building.

Under 10 kW/rack
Legacy enterprise. Raised floor and CRAC units work fine. Fix air management before anything else.
10–40 kW/rack
Mainstream cloud. Contained aisles, chilled water, economization. Air still works if it is done properly.
40–100 kW/rack
Rear-door heat exchangers or direct-to-chip. Water reaches the rack. The hall needs piping.
100 kW+/rack
Direct-to-chip is the default. Power distribution, floor loading, and the building all change with it.

The water-versus-electricity trade

Evaporating water is a cheap way to make cold, so evaporative cooling lowers PUE and raises water use. Rejecting heat dry does the reverse. Neither is right in the abstract; the answer depends on the climate, the local politics of water, and, crucially, on how warm the IT loop can run. That last one is under your control: a direct-to-chip loop that accepts 40 °C water can usually be rejected dry, which removes the water question and much of the compressor load at the same time. Raise the loop temperature first, then choose the rejection method.

Engineering factors

FactorWhy it matters
Rack densityDecides cooling architecture, busway rating, floor loading, and often the building. It is the first number to fix and the hardest to change later.
Loop temperatureThe most useful lever in the whole facility. Warmer loops mean more economizer hours, dry rejection, and lower compressor energy.
Air managementContainment, blanking panels, and sealed cutouts routinely recover 20–40% of cooling energy in a legacy room, for very little money.
Floor loadingA populated liquid-cooled rack can exceed 1,500 kg. Many raised floors cannot take it, and this surprises retrofit projects late.
Redundancy targetTier III concurrent maintainability against Tier IV fault tolerance is a 30–50% capital difference. Choose it from how the application fails over, not from habit.
Ride-throughBatteries give 5–15 minutes, flywheels 15–30 seconds. The shorter one is only acceptable if the generators are genuinely reliable and tested.
Fire and codeLithium batteries, containment, and immersion tanks each change the fire strategy. Involve the authority having jurisdiction before design freeze, not after.
Retrofit disruptionSome upgrades go in rack by rack in a live hall; others need the room emptied. That distinction usually matters more than the capital cost.

Economic and schedule factors

FactorWhy it matters
Power availabilityThe binding constraint on nearly every project. An interconnection date is worth more than a land price, which is why powered land trades at a premium.
Equipment lead timeTransformers, switchgear, and generators all run over a year. Procurement belongs at the front of the schedule.
Stranded capacityFacilities routinely have 10–30% of power or cooling unusable because it is in the wrong place. Instrumentation is what finds it.
Density mismatchA hall designed for 10 kW racks cannot host AI hardware. Existing colocation contracts often lock in a specification the tenant has outgrown.
Water politicsA water number that looks fine in an engineering model can stop a project locally. Reclaimed water and closed loops remove the objection.
Build vs colocateBelow roughly 5–10 MW, colocation is usually faster and cheaper all-in. Above it, self-build wins on unit cost if the demand forecast holds.
UtilizationA half-loaded facility has a worse PUE and worse economics than the same facility full. Phasing the build matters as much as sizing it.

Why the industry average PUE stopped improving

The weighted average PUE reported by Uptime Institute has sat near 1.5 for six years, which surprises people who follow hyperscale announcements of 1.1. Both are true. The easy gains, containment, raised set points, variable-speed fans, and economizers, were captured a decade ago at sites that could take them, and what remains is a long tail of legacy rooms where the building itself is the constraint. New capacity is much better than the average, but it is being added to a fleet that is mostly old. When someone quotes a PUE, ask whether it is one new building or an operator's whole estate.

Core takeaway

Design from the rack outward and from the heat rejection backward, and make them meet. The rack density fixes the cooling architecture; the heat rejection method fixes how much water you spend and how many compressor hours you avoid; the loop temperature is the one variable that improves both at once. Everything else on this sheet is a consequence of those three, and the projects that go wrong are almost always the ones that picked a building first.

Key questions for engineering decisions

Key questions for investment and business analysis

Head-to-head: how do you cool this rack

Rack density is the first fork, so this is the first table. Options are ordered by how much heat they can take out of a cabinet, and the right answer is usually the least invasive one that clears the density you actually need at end of life. The tables after it cover backup power, heat rejection, and how to get the facility at all.

ApproachDensityLoop tempRetrofitPick it when
Air with containmentUp to 20–40 kW18–27 °C supply airDrop-inThe racks fit under 30 kW. Containment is the cheapest capacity in the building and should be done before anything else.
Rear-door heat exchanger20–60 kWChilled waterDrop-in, rack by rackYou need more density in an existing hall and cannot dictate what hardware arrives. The standard colocation answer.
Single-phase direct-to-chip60 kW to well past 10030–45 °CHall retrofitCurrent AI hardware, which ships configured for it. Warm water means the plant can economize or reject dry.
Two-phase direct-to-chipExtreme heat fluxBoiling near chip tempNew buildChip powers past what a single-phase plate can hold. Ask about the fluid's PFAS status before committing.
Single-phase immersion100 kW+ per tankWarm fluidNew buildUniform fleet you can modify, no legacy air plant, and a dusty or humid site. Check server warranties first.
Two-phase immersionHighest availableBoils near 50 °CNew buildThermally the best and commercially parked, because the enabling fluids are PFAS compounds being regulated out.

Carrying the load when the grid drops

Backup is two separate questions: what covers the seconds before the generator picks up, and what covers the hours after. Getting the first wrong is a data loss event; getting the second wrong is a long outage.

OptionRide-throughFootprintReplacement cyclePick it when
Double-conversion UPS with lithium5–15 minutesModerate8–10 yearsThe default. Also gives clean power, and a large plant can bid into demand response between outages.
Double-conversion UPS with lead-acid5–15 minutesLarge3–5 yearsAn existing battery room already sized and suppressed for it, with the replacement cycle already funded.
Flywheel or rotary UPS15–30 secondsSmallAbout 20 yearsGenerators are reliable and tested, batteries are a maintenance burden, and the site runs hot.
Diesel generatorHours to daysLarge yardDecadesStandby duty, essentially always. The constraint is the air permit, not the engine.
On-site gas prime powerContinuousPower plantDecadesInterconnection is years away and the compute cannot wait. Costs more per kWh than the grid.
Fuel cellsContinuousModerateStack every few yearsAir permits block engines. Higher capital, much lower local emissions, still burning gas.

Where the heat finally goes

The last step out of the building is a straight trade between electricity and water, and how much room you have to make that trade depends on how warm the loop runs.

MethodWater useReachesFootprintPick it when
Open cooling towerHighWithin 3–5 °C of wet bulbCompactHot climate, water available and uncontroversial, and the load needs genuinely cold water.
Dry coolerNoneA few degrees above dry bulbLargeThe IT loop runs warm, which liquid cooling allows. Removes the water question entirely.
Adiabatic hybridLowBetween the twoLargeYou want dry operation most of the year without sizing the whole plant for the design day.
Air-side economizerLow to noneOutside air directlyLarge ductsCool clean dry climate. Bring a plan for smoke events and outdoor air quality.
Heat reuse to district networkNoneSells the heatPlant plus connectionA network exists within a few kilometers and the planning authority values it. Not a revenue play.

Build, lease, or rent

Below a certain size, running your own critical facility costs more than it saves. The crossover is mostly about how many megawatts you need and how confident the forecast is.

ModelTime to capacityUnit costFlexibilityPick it when
Hyperscale self-build2–4 yearsLowest at scaleTotal control, large commitmentYou need hundreds of MW, control the hardware, and can carry the balance sheet.
Prefabricated modular6–12 monthsHigher per MWIncrements, vendor-tiedSpeed beats capital cost, or site labor is scarce. Power and cooling skids capture most of the benefit.
Wholesale colocation3–12 monthsMiddleWhole halls, long leases1–10 MW with a stable forecast, and you would rather not operate the building.
Retail colocationWeeksHighest per kWCabinet by cabinetInterconnection density matters, or the requirement is small and uncertain.
EdgeWeeksHighest all-inMany small sitesLatency or data locality genuinely requires proximity. Most workloads do not.