Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    CDU
    Liquid Cooling
    Data Center

    Liquid Cooling Monitoring: CDU, Loop and Distribution Branch

    July 20, 2026
    9 min read

    A CDU, liquid cooling loop, and distribution branch should be monitored as one connected operating path. The minimum useful view includes equipment state, communication health, supply and return temperature, flow, pressure differential, leak detection, trend history, alarms, and the relationship between each branch and the racks it cools.

    The individual metrics matter, and so do the relationships between them. When flow drops on one branch, operators need to know which racks depend on it. When a CDU collection link fails, they need to know whether the equipment failed or only the monitoring path failed.

    What should be monitored on a CDU?

    Monitor the CDU as both a physical cooling unit and a data source. The physical operating view should include the states exposed by the installed CDU. The source material lists unit operating status, alarm status, temperature data, flow related data, pressure data, and branch relationships. The data collection view should include communication status, authentication status, the last successful collection, and a fallback collection path where one is available.

    This distinction is important. A CDU can be cooling normally while the monitoring platform loses authentication. In that case the platform should not report the unit as failed; it should report that CDU telemetry is unavailable. The source example specifically includes a CDU collection interruption and a local PLC fallback path, which is a useful pattern for resilient monitoring.

    What temperatures should be monitored?

    Monitor supply and return water temperatures at the points exposed by the cooling design. The source material specifically uses inlet and outlet water temperature, and these values should be trended over time. A single current reading answers "What is the temperature now?" A trend answers "Is it drifting, oscillating, or changing with workload?"

    Compare temperatures with flow, pressure differential, rack power, cooling alarms, and peer branches.

    Do not create one universal temperature threshold for every cooling system. The acceptable range depends on the installed equipment and design, so use the engineering limits for the specific environment. The monitoring platform should store those limits and show when a measurement approaches or crosses them.

    Why should supply and return temperatures be shown together?

    Showing supply and return together makes it easier to understand how the cooling path behaves under load. If supply temperature remains stable and return temperature rises as compute load rises, the change may reflect the heat being removed. If both temperatures rise unexpectedly, the issue may be upstream. If return temperature rises while flow is falling, the branch deserves investigation.

    The exact diagnosis depends on the system, but the paired view gives operators context. That is why the source interface combines multiple thermal and pressure signals into one 24 hour trend view.

    What flow data should be monitored?

    Monitor current flow and the trend for each branch or loop where the device exposes it. Flow is important because liquid cooling depends on moving coolant through the heat removal path.

    The source example includes branch specific flow warnings, which means the monitoring model should go below total CDU flow. A total value can look normal while one distribution branch is underperforming. For each branch, show the current flow, the expected operating range, the warning state, the recent minimum and maximum, the trend, and the related racks.

    If the system supports several branches with different design capacities, use branch specific thresholds. A small branch and a large branch should not share one arbitrary alarm value.

    What pressure should be monitored?

    Monitor pressure or pressure differential at the points exposed by the cooling equipment. The source design specifically tracks pressure differential trends together with water temperatures, and pressure behavior gives another view into the condition of the cooling path. As with temperature, you interpret it against the approved design range and the historical pattern.

    Look for sudden change, persistent drift, differences between similar branches, and correlation with flow changes or alarms. Do not convert every small movement into an incident. The monitoring system should distinguish normal variation from sustained abnormal behavior.

    What should be monitored on a distribution branch?

    A distribution branch should be treated as a first class managed object. Give it an identity, map it to the CDU or loop that supplies it and to the racks it serves, and collect the branch telemetry.

    A useful branch record includes:

    • Branch name or ID
    • Parent CDU
    • Supply temperature
    • Return temperature
    • Flow
    • Pressure differential
    • Valve state where exposed
    • Leak detection zone
    • Alarm state
    • Collection state
    • Associated racks

    This makes the branch usable in incident analysis and capacity planning. Without an identity and relationships, a branch is only a group of sensor values.

    Why should valve state be included?

    Valve state helps explain whether flow behavior matches the physical control state. If the system exposes valve position or open and closed state, collect it. That becomes especially important when leak workflows can trigger valve closure.

    The monitoring record can then show a complete operational event: the leak is detected, the valve command is authorized, the valve closes, branch flow drops, the affected racks are isolated, and a work order is created. The source material requires linked actions such as valve closure to be authorized and audited, and monitoring the valve state confirms that the requested physical action actually happened.

    How should leak detection be represented?

    Leak detection should be represented as a physical sensor zone linked to a branch or rack area. The source interface uses multiple leak detection lines distributed by zone.

    A leak event should identify the sensor or detection line, the physical zone, the associated branch and racks, the time, the current status, and the response workflow. That works better than one global "water leak" alarm. The response team needs a location, the compute operations team needs an impact list, the facilities team needs the relevant valve or branch, and the workflow system needs the responsible owner. One relationship model can connect all of those.

    What should be monitored in the collection path?

    Monitor whether the monitoring system itself is successfully collecting the data. This includes authentication failures, timeouts, communication interruption, stale data, the last collection timestamp, and fallback source status.

    The source collection architecture explicitly says collection failures should never be silent. That is particularly important for liquid cooling, because an empty chart can otherwise be mistaken for a stable condition. If a CDU stopped reporting ten minutes ago, the dashboard should show "data unavailable" instead of reusing the last value without warning. Staleness should be visible.

    How long should trend history be displayed?

    The source interface uses a 24 hour trend view for temperature and pressure. That is a useful operational window because it shows the current daily behavior and recent changes. Longer history can support maintenance and capacity analysis, and the exact retention period depends on the operations requirement.

    A practical approach keeps recent high resolution data for incident analysis, daily and weekly trend views for operations, and longer aggregated history for planning. What matters is keeping enough history to distinguish one short event from a developing pattern.

    How should liquid cooling alarms be structured?

    Use alarm rules that preserve both the metric and the physical context. An alarm should identify:

    • CDU or branch
    • Metric
    • Observed value
    • Threshold or rule
    • Duration
    • Related racks
    • Severity
    • Collection source
    • Recommended response

    Avoid alarms that say only "temperature high." The operator should not have to search another system to find where the sensor is. The source monitoring design also connects anomalies to work order creation, which makes the alarm part of a response process instead of the end product.

    How should racks be connected to cooling branches?

    Maintain an explicit relationship from branch to rack. This relationship supports three activities. For incident impact, if a branch fails, it lists the racks at risk. For capacity planning, if a branch has little remaining cooling headroom, it tells you not to deploy another high load server into its racks. For maintenance, if a branch is scheduled for work, it identifies the compute resources that may need protection or rescheduling.

    This relationship should live alongside other infrastructure dependencies. For the wider monitoring chain, how liquid cooling monitoring works in high density data centers explains how CDU, branch, rack, and workload context fit together.

    How should liquid cooling data connect to CMDB?

    CDU units, branches, racks, servers, and workloads should share stable identities and relationships. The CMDB does not need to store every second of temperature data, since the time series system can do that. The CMDB should store the objects and relationships needed to interpret the telemetry.

    For example, CDU 02 supplies Branch 07, and Branch 07 cools Rack R18. Rack R18 contains Servers S21 to S28, Server S24 hosts GPU Node N24, and Node N24 runs Training Job J104. With that chain in place, a Branch 07 flow alarm can be traced to the active workload.

    The article on how a CMDB can connect servers, GPUs, containers, applications, business services, and owners describes that relationship pattern more broadly.

    How should thresholds be managed?

    Thresholds should come from the equipment design, commissioning baseline, and operating experience. Do not use one default threshold for every branch. A branch serving four racks may have a different expected flow from one serving one rack, and a new high density deployment may change the normal operating range.

    The platform should therefore support per device and per branch thresholds, warning and critical levels, persistence windows, trend based rules, and baseline comparison. Tune the rules after real incidents. The objective is to identify meaningful degradation without creating constant alarm noise.

    What should a good CDU dashboard show first?

    The first screen should answer whether the cooling chain is healthy and where attention is required. Show CDU status, branch status, branches with warnings, leak state, collection failures, current temperature and flow exceptions, the recent pressure trend, and affected racks. Then allow drill down into individual trends and device details.

    A platform example that combines CDU, branch, leak, and rack relationships is Sensaka.

    If I had to reduce liquid cooling monitoring to one requirement, it would be this: every abnormal reading must be tied to a named physical path and a known set of racks. Once that relationship exists, the metrics become actionable instead of being isolated facilities telemetry.

    Frequently Asked Questions

    What are the most important CDU monitoring points?

    Start with operating state, alarm state, communication status, supply and return temperature, flow, pressure or pressure differential where exposed, and the relationship to downstream branches.

    What should be monitored on each liquid cooling branch?

    Track branch identity, flow, supply and return temperatures, pressure differential, valve state where available, leak status, alarms, and the racks or devices that branch supplies.

    Why should collection health be monitored too?

    If the platform only reacts to equipment alarms, missing telemetry can look like a healthy system. Monitoring the collection path makes authentication failures, timeouts, and communication interruptions visible.