Why the seconds after failure matter to data centre cooling
Authors
James Wild
View bioDuring a power failure, cooling systems can stop almost instantly while IT equipment continues generating heat. This causes rapid temperature increases that can lead to equipment throttling or shutting down within seconds unless resilience measures such as UPS-backed pumps or thermal storage are in place.
As data centres continue to increase in density, designing for cooling resilience is becoming more challenging.
For many facilities - particularly those supporting artificial intelligence (AI) workloads and liquid cooling to servers - the challenge for operators goes beyond achieving cooling capacity during normal operation. Understanding what happens in the moment of failure is vital and needs to be considered in the design.
A power failure may last only a few seconds before backup systems kick in, but that time matters. The heat-rejection plant is affected and will stop cooling immediately, while UPS-backed IT equipment will continue to generate significant heat. Within moments, this can cause IT equipment to throttle or shut down and lead to rapid temperature increases throughout the chilled water network.
With uptime being a key factor in data centres' commercial success, having systems shut down when they get too hot poses a huge risk to commercial viability. In addition, high temperatures can stress equipment, leading to further costs down the line.
Compared with traditional air-cooled facilities, liquid-cooled systems must handle far greater heat loads, so they tend to have lower thermal inertia, posing an additional challenge. Operators increasingly need to understand how long chilled and technical water system remain within acceptable limits during a transient event.
This is where transient modelling becomes valuable.
Why is transient analysis becoming more important?
Traditional cooling calculations are typically based on stable operating conditions. They assess how a system will perform under a steady load once temperature and flow rates have settled and balanced.
During a utility outage, the chilled water system changes operation. Pumps can transfer to UPS power if available, and chillers can restart under generators. As a result, the cooling system is interrupted, and water temperatures rise throughout the system.
In high-density data centres, this happens in a matter of seconds.
Liquid cooling systems tend to operate within tighter thermal boundaries. This means that even relatively short temperature increases can become significant to operation. As rack densities continue to increase, resilience strategies that were acceptable in the past may prove inadequate in the face of failure. We have seen the risks of this in Sydney, where such high return water temperatures can make it difficult to restart chillers after a failure. By the time the chillers restart, the water temperature is so high that it can cause overpressure in the evaporators, causing the chillers to shut down again even after a few minutes.
This is driving the greater use of transient modelling during design stages. Rather than reviewing a single operating condition, engineers can simulate second-by-second behaviour in failure scenarios to understand how temperature, flow, and cooling capacity change over short time intervals.
At Cundall, this analysis is carried out using our in-house H20SAFE transient modelling software. This has been specifically developed to assess the performance of chilled water systems during failure events.
Understanding thermal resilience is key during an outage
One of the key components, often assessed through transient modelling, is the thermal energy store, a tank designed to temporarily hold chilled water.
Its role is relatively simple in principle. During a mains power failure, the thermal store provides a temporary reserve of chilled water while chillers are offline or restarting. That stored thermal energy slows the rate of increase in the system's supply temperature. This buys more time while the mechanical cooling system is restored.
The effectiveness of this depends heavily on the system configuration.
Recently, we undertook modelling of chilled water system performance under a total mains utility loss, considering both a high-density, liquid cooling-enabled arrangement and a lower-density design. Multiple configurations were assessed, including systems with and without mechanical UPS support, and at varying thermal storage volumes.
The team focused their simulations on the entire chilled-water network. This consists of chillers, pumps, fan-wall units (FWUs), coolant distribution units (CDUs), and thermal stores. This allowed us to predict the server and cooling water temperatures throughout the failure period.
The findings showed that without thermal buffering, temperatures can rise quickly. In one scenario without mechanical UPS support or buffer vessels, the model predicted that the ASHRAE (American Society of Heating, Refrigerating and Air-Conditioning Engineers)-defined temperature thresholds would be breached within seconds.
Adding UPS support to distribution pumps and FWUs improved system response, delaying temperature increases and slightly reducing breach durations. However, the most significant improvement came from the introduction of thermal storage.
The modelling allowed us to verify that the thermal storage volume proposed was sufficient to maintain the server inlet temperatures within the client’s required limits for the high-density configuration. In the liquid cooling system we analysed, the worst-case predicted inlet temperature to all CDUs remained within acceptable limits throughout the failure.
Our findings also showed that thermal behaviour varies across different areas of the system. Chiller strings located closest to the load responded differently to those positioned centrally within the network ring. They received and transmitted warmer water through the system more rapidly during these failure conditions.
How we validated the results
To check that the model reflected real system behaviour, we compared the simulation against data captured during a Level 4 test script completed in October 2025.
The focus was on the thermal stores, and simulated inlet and outlet chilled water temperatures were compared against measured temperatures taken during the live test following a mains power loss scenario.
Our simulations replicated the exact chilled water system configuration to ensure an accurate comparison. The simulated temperatures were then compared against the real-world results.
The difference between the modelled and measured temperatures was within 0.2°C.
That level of correlation provides confidence to operators and developers who are constantly looking for resilience measures to keep their facilities operational.
Designing for dynamic cooling environments
As data centres move towards high-density computing and tighter thermal conditions, understanding transient cooling is vital.
The industry has focused for years on electrical resilience; however, cooling resilience now demands the same attention. The focus for mechanical engineers isn’t just on how the cooling system responds whilst operational, but on how it behaves in the critical seconds after a power outage. It is during this time that the design choices determine if the facility can remain within its operational temperature limits.
If you are interested in conducting a transient analysis of your reference design, contact a member of our team to learn how we can support you.