VPS.TC
| $
Server Status
Turkey Istanbul, Türkiye
Active
USA New York, USA
Active
Cart Total:
View Cart
What Is Liquid Cooling in Data Centers?
Datacenter

What Is Liquid Cooling in Data Centers?

Avatar of Defne Defne 17 min read 0 Comments
Share:

Quick Summary – Data Center Liquid Cooling

Liquid cooling moves heat from dense server components into a controlled coolant loop instead of asking room air to carry everything. The right design depends on rack density, hardware compatibility, facility capacity, and service skills.

  • Why it matters — High CPU and GPU densities can exceed the practical limits of air cooling.
  • Main designs — Direct-to-chip, single-phase immersion, two-phase immersion, and rear-door heat exchangers solve different infrastructure problems.
  • Measure the whole system — Include pumps, CDUs, heat exchangers, chillers, fans, controls, and server loads when evaluating energy use.
  • Monitor the path — Track flow, pressure, supply and return temperatures, pump state, leaks, and component sensors together.
  • Plan for failure — Isolation valves, redundant pumps, tested alarms, approved fluids, and controlled shutdown procedures are essential.
  • Start with a pilot — Use a measured pilot rack before changing the cooling design across the facility.

Data center liquid cooling removes heat from CPUs, GPUs, or entire servers with a liquid loop instead of asking room air to carry everything away. Cold plates, immersion tanks, and rear-door heat exchangers fit different workloads, but all require measured capacity, compatible hardware, leak detection, and a tested failure plan.

How liquid cooling moves heat in a data center

Room air is a surprisingly expensive way to move a large amount of heat. Data center liquid cooling takes heat from CPUs, GPUs, or whole server assemblies with water or an electrically non-conductive fluid, then carries it to a heat exchanger, coolant distribution unit (CDU), or facility cooling loop.

In a direct-to-chip system, a cold plate sits on the component and transfers its heat to the coolant. In an immersion system, the hardware sits in dielectric fluid. Both approaches reduce the amount of heat that the room air has to carry, but neither removes the need for careful monitoring and maintenance.

🚀 Boost Your Speed with VPS Server!

Speed up your projects with high-performance SSD storage and 99.9% uptime guarantee.

Get VPS Hosting

ASHRAE’s data center guidance focuses on keeping IT equipment within an acceptable operating range rather than chasing the lowest possible component temperature. That distinction matters. A cooler CPU does not prove that the pump, filter, return loop, or leak detector is healthy.

I have not operated an immersion tank in production, so I would not pretend that a few lab measurements make me an immersion specialist. I have, however, spent enough time reading BMC and facility telemetry to know that one reassuring temperature value can hide a problem elsewhere in the cooling path.

From the field

At the hosting company, I learned to compare component temperatures with cooling-path readings. A healthy chip sensor says very little about flow, pump state, or a blocked filter.

Why air cooling reaches a practical limit

Traditional air cooling uses server fans to move hot air into the hot aisle, where the facility removes it. For low- and medium-density systems, this is still practical and easy to service. The trouble starts when several GPUs or high-core-count CPUs share one rack.

☁️ Gain Flexibility with Cloud Server!

Experience the power of cloud with scalable resources and instant backups.

Cloud Server Plans

Nearly all of the electricity consumed by a server eventually becomes heat. A rack drawing 10 kW therefore produces roughly a 10 kW heat load that the facility must remove. The number is simple. Moving that heat through fans, ducts, cooling units, pumps, and heat exchangers is not.

As rack power rises, airflow requirements, fan energy, duct capacity, and cold-aisle design become more demanding. AI training, high-performance computing, rendering, and scientific simulation make this limit easier to see. A VPS customer may never see a coolant hose, but the density and thermal stability of the physical host still affect service continuity.

Liquid cooling is not mainly about making a processor cold. Its useful purpose is to keep dense equipment stable within its permitted range while reducing the amount of heat that room air and server fans must handle.

What I would do – Measure rack power, server temperatures, airflow, and available room-cooling capacity separately before deciding that liquid cooling is necessary.

Example

A 10 kW rack creates approximately a 10 kW heat load. The calculation is simple, but the facility still has to move that heat through fans, pumps, heat exchangers, or chillers.

Four cooling designs you will encounter

Direct-to-chip cooling

A metal cold plate sits on a CPU, GPU, or another high-power component. A pump moves coolant through channels in the plate, and the warmed liquid returns to a CDU or heat exchanger. From there, its heat moves into a secondary facility loop or the outside environment.

This arrangement targets the components that create the most heat. Memory, network cards, storage, and voltage-regulation components may still depend on airflow, so direct-to-chip does not mean that server fans disappear.

That detail is easy to miss when looking at a product diagram. It is also where deployment plans can become unrealistic.

Single-phase immersion cooling

With single-phase immersion cooling, servers or circuit boards are submerged in an electrically non-conductive fluid. The fluid warms without boiling and travels to a heat exchanger, where it releases the heat before returning to the tank.

Fluid selection and material compatibility are central concerns. Cable jackets, seals, coatings, thermal pads, and adhesives may not behave the same way in every dielectric fluid. Service procedures change too: removing a server can mean draining, lifting, cleaning, and managing fluid rather than simply sliding a chassis out of a rack.

Two-phase immersion cooling

In a two-phase system, the fluid boils at the components and becomes vapor. A condenser above the tank turns that vapor back into liquid, which returns to the system.

Lower pump requirements can be useful, but the design brings its own questions about fluid choice, seals, vapor containment, condenser performance, and environmental handling. I would want those answers in writing from the vendor before putting production hardware into a tank.

Rear-door heat exchangers

A rear-door heat exchanger replaces or sits behind a rack’s rear door. Hot exhaust air passes through the panel, and a liquid loop removes its heat before the air returns to the room.

This can be a useful step when a facility wants to keep standard server designs and avoid installing cold plates inside every chassis. The door still adds weight, pipe routing must be planned, and airflow assumptions need to be checked under real load.

Open Compute Project guidance treats liquid cooling as an infrastructure design problem, not as a server accessory. The rack, building loop, controls, alarms, and service procedures have to agree with one another.

What I would do – Consider direct-to-chip or immersion for genuinely dense GPU racks, and evaluate rear-door exchangers for mixed workloads before changing every server.

Tip

Match the cooling method to rack density and service practice, not to the most impressive specification sheet. A simpler air or rear-door design may be the better choice for mixed workloads.

Direct-to-chip and immersion are not interchangeable

Measure Direct-to-chip Immersion cooling
Contact method Liquid flows through channels in a cold plate attached to selected components. Servers or boards are submerged in dielectric fluid.
Air requirement Some airflow is usually still needed for memory, storage, and network components. Fan requirements can fall substantially when the tank and hardware are designed for it.
Service work Cold plates, hoses, quick disconnects, and seals need inspection. Hardware may need to be lifted from a tank, drained, cleaned, and handled with fluid controls.
Hardware compatibility CPU, GPU, chassis, cold plate, hose, and connector compatibility must be confirmed. Cables, seals, coatings, adhesives, thermal materials, and fluid compatibility need broader review.
Ease of transition It can remain relatively close to an existing server design. It can require larger changes to racks, tanks, hardware, and service procedures.

No method wins everywhere. Existing rack constraints, maintenance experience, coolant availability, redundancy, and the building-side water loop all affect the decision.

The risk of fluid contacting electronics is different as well. A direct-to-chip loop is normally closed and controlled. An immersion system exposes much more of the assembly to fluid, so changing a seal, cable jacket, or thermal pad without manufacturer approval is a real compatibility risk.

Start small.

The parts that make a liquid loop dependable

Liquid cooling is more than attaching hoses to a server. A typical direct-to-chip installation includes these parts:

  • Cold plate: A metal assembly that transfers heat from a CPU or GPU into the coolant.
  • CDU: A unit that may separate the facility loop from the IT equipment loop and include pumps, filters, sensors, and controls.
  • Heat exchanger: Equipment that moves heat from the IT loop to building water or another cooling circuit.
  • Quick disconnects and hoses: Components that allow a server loop to be isolated during replacement; connector standards and drip behavior matter.
  • Flow and temperature sensors: Devices that identify low flow, high inlet temperature, and pump problems.
  • Leak detection: Cable sensors, drip trays, or cabinet detection systems that generate an alarm.
  • Control and redundancy: Pump, power, and control paths designed so one failure does not immediately stop cooling.

Water quality belongs in the design document. Conductivity, corrosion, biological growth, and particle buildup need defined limits and monitoring. Facility water is not automatically the same as the fluid approved by the equipment manufacturer. Connecting city water directly to an IT loop is not a safe shortcut.

The electrical side deserves the same care. Pumps and CDUs need properly designed redundant power, and those loads should be monitored like servers. The measurement and phase-balancing approach in What Is a Data Center PDU? Power Distribution Guide applies to cooling equipment too.

What I would do – Record every hose’s source, destination, valve, sensor, and isolation step in the facility documentation. A label on the rack is useful; a tested diagram is better.

What liquid cooling changes in the power budget

Liquid carries heat efficiently, which makes high-density racks easier to support. Pumps and CDUs add their own electrical load, though. Pump power, coolant temperature, flow rate, heat-exchanger efficiency, and outdoor conditions all affect the total.

Comparing only server fan power gives you the wrong picture. I would measure these loads together:

  • Server and GPU load
  • Server fan consumption
  • CDU pump consumption
  • Heat exchanger and dry-cooler load
  • Chiller or cooling-tower consumption
  • Controls, monitoring, and auxiliary equipment

Power usage effectiveness (PUE) includes energy used outside the IT equipment, so liquid cooling does not automatically improve it. Compare the design with the site’s climate, water temperature, free-cooling opportunities, and expected operating load. Fluid replacement, filters, maintenance, training, and spare parts belong in the calculation too.

The same measurement discipline applies when comparing owned capacity with cloud capacity. A site carrying a steady GPU load may benefit from owning the infrastructure; occasional workloads may still be cheaper to outsource. How to Reduce Cloud Costs: 10 Practical Optimization Tips uses the same resource-measurement principle.

Do not hide pump and CDU consumption in a separate spreadsheet.

Monitoring has to follow the heat path

Server temperature alone is not enough. At minimum, track supply and return coolant temperature, flow rate, pressure, pump state, CDU alarms, leak sensors, and component temperatures.

On Linux servers, hardware sensors are a useful first check:

sudo apt install lm-sensors
sudo sensors-detect
sensors

These commands list sensors visible to the operating system. Seeing CPU temperature output does not prove that the CDU or leak detector is being monitored. Those values often arrive separately through a BMC, SNMP, Modbus, or vendor API.

For a BMC, I might start with:

ipmitool sensor | grep -Ei 'temp|fan|flow|pump|leak'

This is only a filter. BMC sensor names vary, so an empty result is not automatically a failure. Check the vendor’s sensor map and thresholds before changing anything. Record normal operating values first; otherwise, a perfectly ordinary startup condition can become a false alarm.

I normally use three alarm levels: informational, intervention required, and controlled shutdown. If flow drops, reduce the thermal load and isolate the affected loop according to the facility procedure. For process load, Linux Process Management: Using ps, top, and kill is useful, but process usage cannot replace liquid-loop telemetry.

SSD and disk temperatures deserve attention too. The device differences explained in What Is an SSD? Full Form, Speed and Benefits matter here: a normal CPU cold-plate temperature does not prove that an SSD is within its own limits.

What I would do – Put coolant temperature, flow, pressure, pump state, and component temperature in the same event record. One graph rarely tells the whole story.

Caution

An empty BMC sensor query does not prove that a pump or leak detector is broken. Sensor names differ between vendors, so check the documented sensor map before changing thresholds.

Maintenance, leaks, and controlled failure

The maintenance plan should name the manufacturer-approved fluid, filter interval, connection checks, sensor tests, and pump inspections. A closed loop still needs attention. Small leaks, loose connectors, or clogged filters can reduce performance long before someone sees water on the floor.

For every rack, I want clear answers to these questions:

  • Which valve isolates only the affected server or rack?
  • How long after flow stops should the alarm arrive?
  • At what temperature should virtual machines move to other hosts?
  • In what order are electrical and liquid isolation performed after a leak?
  • How is the loop drained or sealed before a server is removed?
  • Where are spare pumps, hoses, connectors, and approved fluid stored?

A failure procedure is not trustworthy until it has been tested under controlled conditions. Put a sensor into a temporary test mode, start the backup pump, and confirm that the alarm reaches the on-call team. Run this during a maintenance window, with a clear rollback plan.

Change management matters here too. Adding a GPU, replacing a cold plate, or increasing flow should be followed by thermal measurements. Exceeding the specified pressure, connector type, or fluid compatibility limit for a small temperature improvement is a poor trade.

I learned this habit from less exotic systems: if a recovery procedure exists only in a document, I assume it will fail at the worst possible time. Test it while everyone is calm.

What I would do – Test pump-stop, low-flow, and leak scenarios, then measure the time from alarm to controlled shutdown.

When liquid cooling earns its place

Liquid cooling is not the first choice for every physical server. A well-designed air system can be simpler for medium-density web, database, or file servers. The case for liquid grows with high-density GPU racks, limited room-cooling capacity, or a new facility designed around high-power compute.

Use these questions to frame the decision:

Question Data to evaluate
How much heat does the rack produce? Average and peak kW, GPU count, and CPU load.
What can the building support? Cooling capacity, piping, floor load, and electrical redundancy.
Is the service team ready? Fluid handling, safe removal, alarms, and emergency training.
Is the hardware compatible? Cold plate, chassis, seals, warranty, and manufacturer approval.
What is the cost? Capital cost, energy, maintenance, fluid, spare parts, and operating cost per unit of capacity.

The safest route is a small pilot rack. Measure more than peak-load temperature: include low-load pump behavior, flow after a reboot, maintenance time, alarm delay, and serviceability. A design that looks efficient on paper changes the cost calculation if every maintenance task takes hours.

The aim is not to say that a facility uses liquid. The aim is to manage a defined power density safely, measurably, and in a way technicians can maintain. If that limit is unclear, measure capacity first and choose the cooling method second.

A migration path I would trust

  1. Collect rack power, temperature, fan, and room-cooling data for several weeks.
  2. Select the highest-density rack and measure the real limit of its air-cooling design.
  3. Compare direct-to-chip, rear-door heat exchangers, and immersion with their maintenance requirements.
  4. Verify manufacturer compatibility, warranty conditions, and the technical data sheet for the proposed fluid.
  5. Prepare redundancy plans for CDUs, pumps, sensors, and alarm components.
  6. Build a pilot rack and test normal, peak-load, reboot, and failure scenarios.
  7. Evaluate the result against capacity, energy, maintenance time, and incident-response cost.

This approach works in a small hosting facility as well as in a large compute cluster: make the load and the limit visible first, then select the cooling method. Liquid is not a repair for infrastructure nobody has measured.

What I would do – Keep the pilot’s baseline, alarm history, maintenance notes, and energy readings together. When the next rack is planned, you should be able to show what happened, not just what the vendor promised.

Before You Deploy Liquid Cooling

  • Measure average and peak power for each candidate rack.
  • Confirm CPU, GPU, chassis, seal, cable, and warranty compatibility.
  • Document the approved coolant and its required water-quality limits.
  • Design isolation valves and redundant power for pumps and CDUs.
  • Install and test flow, pressure, temperature, and leak sensors.
  • Write the alarm, workload-migration, and controlled-shutdown procedures.
  • Run a pilot through normal, peak-load, reboot, and failure scenarios.

If your racks are approaching the limits of air cooling, start by measuring power, airflow, and thermal behavior for one pilot rack. The useful decision is not whether liquid cooling sounds advanced, but whether it makes your specific workload easier to run and service.

Explore VPS plans

Frequently Asked Questions

What is data center liquid cooling?

Data center liquid cooling removes heat from server components with a liquid rather than relying only on room air. In direct-to-chip systems, coolant flows through cold plates attached to CPUs or GPUs. In immersion systems, servers or boards sit in dielectric fluid. Rear-door heat exchangers cool hot exhaust air after it leaves the server.

Is liquid cooling better than air cooling?

It is better for some density and capacity problems, not automatically for every server. Air cooling is often simpler for low- and medium-density workloads. Liquid becomes more attractive when GPU or CPU heat exceeds practical airflow limits, room-cooling capacity is constrained, or a new facility can support liquid loops and their maintenance requirements.

Does direct-to-chip cooling remove the need for fans?

Usually not completely. Cold plates can remove heat from CPUs and GPUs, while memory, storage, network cards, and other components may still need airflow. Fan requirements can be reduced, but the result depends on the server design and the components covered by the liquid loop.

Is liquid cooling safe for electronic components?

It can be safe when the system uses approved hardware, compatible fluids, suitable seals, controlled pressure, and leak detection. Direct-to-chip systems usually keep liquid in a closed loop. Immersion systems expose more of the assembly to fluid, so cable jackets, coatings, connectors, and manufacturer compatibility need close review.

What should be monitored in a liquid cooling system?

Monitor supply and return temperatures, flow rate, pressure, pump state, CDU alarms, leak sensors, and server component temperatures. Record normal values before setting alarm thresholds. BMC or operating-system sensors may not include facility equipment, so collect CDU and building-side telemetry through the appropriate interface.

How should a facility start using liquid cooling?

Begin with a measured pilot rack rather than changing the entire room. Record rack power, temperatures, pump behavior, flow after reboots, alarm delays, maintenance time, and failure responses. Confirm hardware and fluid compatibility, test isolation and controlled shutdown, then compare the pilot's energy, capacity, and service costs with the existing air-cooled design.

Avatar of Defne
Author

Defne