The Physics Air Cooling Can't Argue With
A standard data-center air-cooling system works by moving large volumes of cold air past hot components and exhausting the warmed air into a return plenum. For years this was perfectly adequate: a typical compute rack drew five to ten kilowatts, and a well-designed raised-floor or hot-aisle/cold-aisle layout could handle that without heroics. Modern AI racks running dense GPU nodes have torn up that calculus entirely. A single rack of current-generation accelerators routinely hits 40 to 80 kilowatts; some configurations push past 100 kW. Air's fundamental problem is its low heat capacity — to remove 100 kW via air alone you need airflows so violent they become acoustically and mechanically impractical, and the downstream cooling plant needed to handle the hot exhaust becomes enormous and expensive. Water carries roughly 3,500 times more heat per unit volume than air. That gap is why liquid cooling has moved from HPC curiosity to mainstream datacenter requirement inside a few years.
Modern AI racks running dense GPU nodes have torn up that calculus entirely.
The Ladder of Liquid Options
Operators don't face a binary choice between air and full immersion. There's a practical ladder of interventions, each step adding cooling capacity and engineering complexity.
Rear-door and close-coupled heat exchangers sit at the bottom rung. A rear-door heat exchanger (RDHx) replaces the standard rack door with a panel threaded with chilled-water coils; air still flows through the servers normally, but it's cooled again before leaving the rack. Close-coupled in-row units park cooling capacity directly beside the racks rather than pushing chilled air from a distant CRAC unit. Both approaches are retrofit-friendly — existing servers need no modification — and they extend the practical ceiling for air-cooled racks meaningfully, though they can't keep pace with the highest AI densities. Think of them as a bridge technology: valuable in mixed-density environments where some racks are conventional and a few are hot.
Direct-to-chip liquid cooling is where serious AI clusters live today. Cold plates — metal blocks with internal channels — are bolted directly onto the processors and other high-heat components. A closed loop carries a coolant (typically water-glycol) from a facility manifold, through the cold plate, and back to a heat exchanger or cooling distribution unit (CDU) outside the rack. The GPU or accelerator dumps its heat directly into the fluid rather than into the surrounding air, which means dramatically less airflow is needed inside the chassis. Some vendors now ship servers where the liquid circuit handles 70 to 80 percent of total rack heat, with residual air cooling managing memory and ancillary components. This approach integrates cleanly with standard server form factors and allows normal maintenance — a failed node can be swapped after disconnecting quick-release liquid couplings. The facility side requires new wet infrastructure: supply and return manifolds, leak-detection systems, and careful management of dew point to prevent condensation. It's a significant change to facility design, but one the major colocation and hyperscale operators have now largely committed to.
Full immersion cooling takes the logic to its endpoint: the entire server is submerged in a tank of dielectric fluid — either a mineral oil derivative or a purpose-engineered fluorocarbon fluid — that conducts heat away from every component simultaneously. Single-phase systems circulate the fluid continuously; two-phase systems exploit the latent heat of the fluid's evaporation and recondensation cycle for even higher efficiency, at the cost of more complex fluid management. PUE figures for immersion installations can approach 1.03 to 1.05, genuinely impressive numbers. The friction is real, though. Standard servers aren't immersion-ready — some components, particularly spinning storage and certain plastics, degrade in some fluids. Serviceability changes completely: pulling a node means draining or handling fluid, and every technician interaction becomes a messier proposition. The tank footprint is also large and inflexible. Immersion suits greenfield deployments designed around it from the start; retrofitting it into an existing hall is rarely straightforward.
| Rear-door / close-coupled | Direct-to-chip | Immersion | |
|---|---|---|---|
| Capacity | to ~30–40 kW | to ~100 kW+ | beyond 100 kW |
| What changes | heat exchanger on the rack | cold plates, manifolds, a CDU | tanks of dielectric fluid |
| Serviceability | familiar | quick-disconnects, leak detection | fluid handling on every touch |
| Best fit | mixed-density retrofits | the AI-era default | the density frontier |

Where the Threshold Lands
The industry has effectively settled on a rough rule of thumb: below 40 kW per rack, enhanced air or rear-door solutions remain competitive; between 40 and 100 kW, direct-to-chip liquid cooling is the primary architecture; above 100 kW, immersion — or a direct-to-chip system with immersion handling the residual — becomes the serious option. As accelerator roadmaps continue to push per-chip TDP upward, these thresholds keep shifting left. The reign of pure air cooling in AI facilities isn't ending — it has already ended.
See the thresholds cross live: the AI Rack Builder on the homepage walks this exact ladder as you fill a rack.
