AI drives cooling systems to become a core challenge in data center design

As AI Workloads Push Rack Densities into the Megawatt Era, Cooling Is No Longer Just Supporting Infrastructure—It Is the Core of the Entire Data Center Architecture

This theme emerged from the Data Center World conference in Washington, D.C., where speakers included Phil Lawson-Shanks, Chief Innovation and Technology Officer at Aligned Data Centers, and Mauro Atalla, Senior Vice President and Chief Technology and Sustainability Officer at Trane Technologies.

In the past, cooling was simply a facility-level issue—how to remove heat from the white space. Today, it has evolved into a system-level design challenge spanning chips, fluids, controls, and workload scheduling.

tp52

From Heat Removal to System Design

Lawson-Shanks framed the shift directly: heat has always been a constraint in data center design, but AI is changing both the magnitude and behavior of heat.

“We are moving from tens of kilowatts per rack to hundreds of kilowatts, and even toward megawatt-scale racks,” he said. “That fundamentally changes how facilities are designed and operated.”

The old model—deploy infrastructure, hand over the data hall, and respond reactively to load changes—is breaking down. AI infrastructure demands much tighter coordination between IT systems and facility systems.

That includes deep integration with building management systems and data center infrastructure management (DCIM) platforms, and ultimately visibility and control over workload schedulers.

“You have to treat the entire building as one organic, integrated system,” Lawson-Shanks said, “rather than a series of separate layers reacting to one another.”

The Collision of Surging Demand and Technology Transition

Atalla noted that demand growth and technological change are colliding head-on. Operators are scaling toward multi-gigawatt campuses while simultaneously transitioning from air cooling to liquid and hybrid cooling architectures.

“Massive demand, delivery-time pressure, and technological evolution are all converging at once,” he said. “It is like retrofitting a system while it is running.”

Cooling now accounts for roughly 20% of data center energy consumption, making it one of the few adjustable levers for both scaling compute and reducing total power consumption.

That is pushing vendors to start at the chip level and move toward system-level solutions.

“The heat source dictates everything,” Atalla said. “We spend as much time working with chip developers as we do with operators.”

Hybrid Cooling Becomes the Mainstream Choice

Both executives agreed that hybrid air-and-liquid cooling architectures will be the mainstream approach in the near term.

Single-phase liquid cooling is expected to support the thermal needs of several future chip generations. Two-phase cooling systems are under development but remain constrained by complexity and refrigerant challenges.

“Single-phase liquid cooling will eventually fall short of demand,” Atalla said, “but the transition to two-phase must be driven by chip requirements.”

Lawson-Shanks said operators are already designing ahead for this uncertainty.

“The liquid cooling loop has always been a part of the system,” he said. “The key now is how to scale it in coordination with air cooling and respond flexibly as densities change.”

That flexibility extends to facility design itself. Aligned is deploying modular skid-based architectures that package racks, power, and cooling into replicable standard units.

“You can plug in a 2- to 3-megawatt unit and then keep scaling from there,” Lawson-Shanks said.

Cooling Response Speed Becomes Key to Reliability

The shift to liquid cooling also makes time an increasingly important operational challenge. Air-cooled environments can tolerate several minutes of interruption before reaching thermal limits, while liquid-cooled systems can only withstand seconds.

“Liquid cooling has no buffer,” Lawson-Shanks said. “The mechanical side must operate continuously and instantly.”

This is driving new demand for thermal buffering, including thermal storage tanks and tighter integration with electrical systems.

It is also changing the risk-sharing model between operators and customers, especially in defining responsibility boundaries at the rack level. As system integration increases, the question has shifted from “can tighter control be achieved” to “who owns that control.”

Atalla said the technical barriers are largely solved.

“This is not a technology limitation,” he said at Data Center World. “It is a matter of design philosophy and risk tolerance.”

Operators are already aggregating telemetry into centralized data platforms and giving customers partial visibility. But true bidirectional control—where workloads feed back to the cooling system in real time—remains very limited.

“If we can sense workloads coming in advance, we can prepare the cooling state in advance,” Lawson-Shanks said. “But very few environments can do this today.”

Sustainability and Efficiency Goals Are Converging

Despite rising complexity, both executives believe sustainability and performance are converging rather than being in opposition.

“Reducing energy consumption creates cascading benefits,” Atalla said. “There is no trade-off between performance and sustainability.”

This trend is driving interest in waste heat reuse, but implementation varies widely by region.

Europe already has cases of using data center waste heat for district heating and greenhouse heating. In the U.S., however, the geographic distance between data centers and heat users limits similar opportunities.

Nevertheless, Atalla still positions data centers as “thermal energy production plants” that could eventually integrate into the broader energy ecosystem.

“There is no one-size-fits-all solution,” Atalla said. “Over time, the industry will move toward standardization, but right now every organization is trying a slightly different path.”

Q&A

Q1: Why has data center cooling become a core challenge in the AI era?

A: As AI workloads continue to grow, rack densities have moved from tens of kilowatts in the past to hundreds of kilowatts and even megawatt scale. The magnitude and behavior of heat that cooling systems must handle have fundamentally changed. Cooling is no longer just a facility issue; it is a system-level design challenge involving chips, fluids, controls, and workload scheduling. At the same time, cooling accounts for about 20% of total data center energy consumption, giving it dual significance for both scaling compute and reducing power consumption.

Q2: What is the difference in reliability between liquid cooling and air cooling systems?

A: Air-cooled systems can tolerate several minutes of interruption before reaching thermal limits, while liquid-cooled systems, lacking thermal buffering, can only withstand seconds of interruption. This means liquid cooling places extremely high demands on mechanical-side continuity, requires thermal buffering such as thermal storage tanks, and needs tighter integration with electrical systems. It also changes how operators and customers define responsibility boundaries at the rack level.

Q3: How far has data center waste heat reuse developed?

A: Implementation of waste heat reuse varies by region. Europe already has real cases of using data center waste heat for district heating and greenhouse heating, while in the U.S., opportunities are limited by the greater geographic distance between data centers and heat users. The industry views data centers as “thermal energy production plants” that could eventually integrate into the broader energy ecosystem, but there is not yet a unified standardized approach.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top