Nearly all electrical power consumed by servers becomes heat that must leave the data hall. Cooling is the controlled chain that moves heat from chips to room air or liquid, through facility equipment, and finally to outdoor air, water, or another heat sink. The correct design depends on rack density, server environmental limits, climate, water availability, required redundancy, maintenance strategy, and workload behavior.
Cooling choices trade resources rather than offering a universally superior answer. Evaporative systems can reduce compressor energy in suitable conditions but consume water. Air-cooled chillers can reduce onsite water use while drawing more electricity under some weather and load conditions. Direct-to-chip liquid can handle dense racks efficiently, yet it still needs a facility heat-rejection loop and introduces connectors, pumps, controls, chemistry, and leak-response procedures.
What you will learn
- Trace heat from processors through air or liquid to final rejection
- Calculate thermal load and interpret PUE and water tradeoffs
- Compare air cooling, direct liquid cooling, and immersion at system level
Heat removal begins at the component
A heat sink and fan can transfer component heat to server airflow. In a room-scale air system, cold aisles deliver supply air to server inlets and hot aisles collect exhaust. Containment reduces mixing, allowing cooling equipment to operate with more predictable temperatures and airflow. Blanking panels, cable openings, fan control, and rack layout can materially affect performance without changing chiller capacity.
Air carries limited heat per unit volume, so very dense racks require large airflow and careful pressure management. Raising allowable temperature difference can reduce airflow, but equipment inlet specifications and hotspots constrain that choice. Thermal sensors should measure actual server inlets and liquid conditions rather than rely only on room averages that can conceal recirculation.
Liquid moves heat closer to the source
Direct-to-chip cold plates transfer heat from selected components into a technology cooling loop. Coolant distribution units can isolate facility water from server coolant while controlling temperature, flow, pressure, and water quality. Some server heat may still enter room air, so a hybrid design can require both liquid and air capacity.
Immersion places compatible electronics in dielectric fluid, moving heat through fluid circulation and heat exchangers. It can remove fans and support density, but equipment compatibility, maintenance workflow, fluid handling, warranties, material behavior, and fire code require careful evaluation. Neither liquid method eliminates heat rejection; it changes the medium and temperatures at which heat travels.
Heat rejection connects climate to operations
Chillers use a refrigeration cycle to move heat to condenser air or water. Cooling towers reject heat partly through evaporation, while dry coolers use outdoor air without routine evaporation. Economizers can reduce compressor work when outdoor conditions allow. Performance changes with temperature, humidity, fouling, approach temperatures, part load, and control settings, so annual models should use local weather rather than one design-day efficiency.
Water use must be measured with system boundaries. Onsite consumption can include cooling-tower evaporation, blowdown, humidification, and treatment, while electricity generation may have offsite water implications that vary by grid mix. Water usage effectiveness is useful when consistently defined, but no single metric captures local scarcity, source quality, discharge constraints, or seasonal operating limits.
Reliability comes from maintainable architecture
Redundancy labels such as N, N+1, or 2N describe installed paths or units but do not prove end-to-end availability. Common pipes, controls, electrical feeds, valves, or water sources can defeat nominal redundancy. Concurrent maintainability requires isolating equipment for service while keeping required loads within temperature limits, including transition and restart behavior.
Commissioning verifies sequences under realistic conditions: pump failure, power transfer, control loss, blocked filters, leak alarms, and changing IT load. Operators need water chemistry, spare parts, sensor calibration, and trained response. A theoretically efficient design can underperform if controls fight each other or maintenance is impractical, making operational capability part of the cooling system rather than an afterthought.
Common misconceptions
“Cooling is a small auxiliary detail after servers and power are selected.”
Thermal architecture constrains rack density, facility capacity, water and electricity use, uptime, equipment choice, floor layout, commissioning, and capital cost.
“Liquid cooling eliminates air cooling and water use.”
Some components may still reject heat to air, and the liquid loop still needs pumps and final heat rejection that may use dry or evaporative equipment.
Risks and limitations
- Unmodeled hotspots can throttle or damage equipment even when average room temperature appears acceptable.
- Water scarcity, quality, discharge, or permit constraints can limit evaporative-system operation and expansion.
- Common piping, controls, or power can create correlated failures despite redundant cooling units.
- Retrofitting liquid cooling can require structural, piping, electrical, warranty, and maintenance changes beyond the rack itself.
Key takeaways
- Electrical IT load becomes approximately equal thermal load that must be transported and rejected.
- Air, direct-to-chip, and immersion systems solve different density and operational problems.
- Liquid cooling moves heat efficiently but does not eliminate facility heat rejection.
- Climate and water context determine annual performance more accurately than one nominal efficiency.
- Commissioning and maintainability are essential parts of cooling reliability.
Primary and further reading
Test your understanding
Score at least 2 out of 3 to complete this lesson. Explanations appear after you submit.