Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Capacity Planning
    Infrastructure Operations
    Forecasting

    Forecasting Compute, Storage, Network, Power and Cooling Capacity

    July 8, 2026
    10 min read

    Organizations can forecast compute, storage, network, power, and cooling capacity by measuring each resource separately, tracking how its usage changes over time (including planned reservations and growth), and then working out which required resource is likely to reach its limit first.

    The source capacity model keeps returning to the same operating rule: a deployment is constrained by the shortest board. Free compute does not help if power is exhausted, free rack space does not help if cooling is constrained, and free storage capacity does not guarantee enough throughput. Capacity forecasting therefore has to cover several dimensions at once.

    Why should capacity be forecast by domain?

    Different infrastructure resources grow at different rates and have different lead times. Compute demand may climb quickly while storage capacity grows steadily. Network requirements may jump when a new cluster is added, power may turn into the physical constraint, and cooling may become the next limit once power is upgraded.

    If these domains are forecast independently and never compared, the organization can still miss the constraint that actually blocks deployment. The source material explicitly includes:

    • Compute capacity forecasting
    • Storage capacity forecasting
    • Rack and power capacity planning
    • Cooling constraints
    • Network and storage bottlenecks
    • Expansion due-date prediction

    The forecast should bring those domain views together into one planning model.

    What should be forecast for compute?

    Compute forecasting should distinguish installed, available, allocated, healthy, and reserved resources. For AI infrastructure, the source examples include total accelerator cards, available cards, allocation level, resource specifications, queue state, fragmentation, capacity-level alerts, and the forecast expansion date.

    A raw card count is not enough. Different accelerator models may not be interchangeable, a degraded card should not count as fully available, and a free card may not match the memory or topology requirement of a queued job. When the environment is heterogeneous, the forecast should work with resource classes or specifications.

    For the allocation layer, how GPU resource pooling and scheduling work in Kubernetes and AI infrastructure explains why usable compute depends on health, resource shape, and scheduling policy.

    What should be forecast for storage?

    Storage forecasting should include both capacity and performance wherever the workload depends on both. The source infrastructure model monitors capacity, throughput, IOPS, and latency, and it includes storage capacity prediction.

    This matters because a storage system can have plenty of free capacity and still lack the performance for another high-throughput workload. For a simple archival workload, capacity may dominate. For AI training, throughput and checkpoint performance can matter just as much, so the forecast has to be aware of the workload.

    Saying "storage is 60 percent free" does not mean there is room for every future workload. Ask whether the planned workload's capacity and performance requirements can both be met.

    What should be forecast for network capacity?

    Network forecasting should consider the connectivity the workload needs. The source model includes the training network, management network, business and inference entry points, cross-data-center dedicated circuits, packet loss, latency, throughput, and port capacity.

    Total bandwidth is only part of network capacity. A new deployment may require specific high-speed ports, a particular topology, inter-data-center connectivity, storage network capacity, or redundant paths, so the capacity model should include the network profile of the planned workload or server.

    In a physical expansion, a rack with free space and power can still be blocked because no network ports are available.

    What should be forecast for power?

    Power forecasting should use current load, peak behavior, approved capacity, reservations, and the site's redundancy policy. The source rack-capacity model includes A and B feeds, circuit headroom, PDU and UPS data, power-density heatmaps, and peak and off-peak power information.

    Historical average power alone is a poor basis for the forecast, because new server generations can behave very differently. The source case material specifically warns that traditional placement assumptions can go wrong when newer infrastructure draws more power than expected. Use measured production behavior to refine the equipment profile over time.

    What should be forecast for cooling?

    Cooling forecasting should reflect the actual cooling path that serves the planned equipment. The source infrastructure model includes air cooling, CDU, distribution branch, supply and return temperature, flow, pressure differential, leak detection, and cooling headroom.

    For liquid-cooled infrastructure, branch or CDU capacity may set the limit. For air-cooled infrastructure, local rack or zone conditions can become limiting before total facility cooling does.

    The physical relationships matter here. A rack can have electrical headroom and still sit on a cooling branch with little capacity left, so the forecast should link rack, equipment, and cooling path.

    What is the weakest-link rule?

    The weakest-link rule says deployable capacity is set by the first required resource that reaches its limit. The source material states this explicitly for power, cooling, and circuit headroom, and the wider capacity-planning source also names network and storage as possible causes of stranded capacity. The same logic therefore applies to the full deployment profile.

    As an example, say compute supports 20 more nodes, rack space supports 18, power supports 12, cooling supports 14, network supports 10, and storage performance supports 16. The current deployable capacity is 10, with network as the first bottleneck. If network is expanded to support 20, power becomes the next bottleneck at 12.

    That tells a planner far more than six separate percentages would.

    Why should forecasting use a deployment profile?

    Raw free resources do not tell you what can actually be deployed. A deployment profile defines what one planned unit requires. For a GPU node, that can include:

    • Rack U height
    • Power requirement
    • Cooling method
    • Network ports
    • Storage requirement
    • Accelerator type
    • Topology requirement

    The forecast then answers how many more units of this exact profile the site can support, which is easier to act on than a figure for how much total capacity remains. Different hardware profiles can give different answers for the same rack or cluster.

    How should reservations be included?

    Keep reserved capacity separate from generally available capacity. The source planning model includes reservation and future expansion, and a project may reserve rack space, power, cooling, network ports, compute, or storage. If the reservation is invisible, the same capacity can be promised twice.

    Forecasting should therefore distinguish installed, allocated, reserved, free, and deployable capacity. That also explains why a site can look underused while new requests still get turned down: some of the capacity is already committed to future work.

    How should growth trends be used?

    Historical consumption trends are one input to the forecast. The source assistant includes capacity prediction and threshold forecasts, and the source website material also discusses forecasting capacity needs and spotting trends that guide infrastructure upgrades.

    The model does not define one universal forecasting algorithm, so the forecast should be open about its method. A simple trend may be enough in a stable environment, while a more complex environment may use scenario-based planning. Either way, show the current level, recent trend, reserved growth, threshold, forecast date, and confidence or scenario.

    If a forecast date depends on volatile demand, do not present it as certain.

    Why should there be several scenarios?

    Infrastructure demand rarely follows one fixed line. A practical planning model can include baseline demand, expected growth, and a higher-growth scenario, and the source capacity-planning material supports this kind of scenario thinking around current, reserved, expected, and future growth.

    Scenarios help with timing decisions. If the baseline reaches the limit in 12 months but the higher-growth case reaches it in 6, the team can compare the expansion lead time with both possibilities, which beats pretending one date is guaranteed.

    How should lead time affect capacity forecasting?

    Capacity planning should forecast the decision date as well as the exhaustion date, because different constraints take different amounts of time to solve. Adding network ports may be fairly quick. Expanding power can require engineering work, cooling upgrades can require construction, and hardware procurement has its own lead time.

    The forecast should therefore show when action needs to start. If power will become limiting in eight months and the upgrade takes ten months, the operational problem already exists. The source model's "expansion due-date prediction" is only useful when it is connected to the time needed to act.

    How can live operations improve the forecast?

    After deployment, compare planned assumptions with real behavior. The source planning material uses production data to improve future placement and capacity decisions.

    Actual server power may differ from what was expected. Storage throughput may drop under concurrent load. GPU utilization may be lower because of data-supply constraints, and a cooling branch may hit its practical limit earlier than planned. The model should update its resource profiles when the evidence supports it, which creates a feedback loop between operations and planning.

    How should stranded capacity appear in the forecast?

    Report stranded capacity separately from truly deployable capacity. The source material explicitly identifies fragmentation and constraints across several domains as reasons nominal capacity can shrink. Typical cases are free U space with no power, free GPUs with the wrong topology, free storage with too little throughput, and free rack power with constrained cooling.

    The forecast should show both the nominal free resource and the constraint blocking it.

    For a detailed approach, how data centers detect stranded capacity caused by power, cooling, network, or storage constraints explains how to classify the stranded portion.

    What should a capacity dashboard show?

    A capacity dashboard built on the source model can show:

    • Current supply
    • Current allocation
    • Available capacity
    • Reserved capacity
    • Utilization trend
    • Threshold warning
    • Forecast threshold date
    • Limiting constraint
    • Second limiting constraint
    • Expansion lead time
    • Stranded capacity

    For AI infrastructure, it can also show resource fragmentation and capacity by accelerator specification. The dashboard should support drill-down by site, rack, cluster, storage pool, or resource class.

    A platform example that combines capacity-level warning, expansion prediction, fragmentation identification, and infrastructure data across domains is Sensaka.

    If I were designing the planning process, I would stop asking "When will we run out of capacity?" and instead ask "For the next planned workload profile, which resource becomes limiting first, when does that happen under each growth scenario, and how long will it take us to remove that constraint?" That is the forecast the operations team can actually use.

    Frequently Asked Questions

    What should infrastructure capacity forecasting include?

    The source material supports forecasting across compute, storage, network, rack, power, cooling, and related operating constraints rather than using one total-capacity number.

    How should organizations identify the next capacity bottleneck?

    Evaluate the planned workload or server profile against every required resource and identify which dimension reaches its approved limit first.

    Why should forecasting include reservations?

    Reserved capacity is not generally available to other work. Forecasting should separate installed, allocated, free, and reserved capacity so the same headroom is not promised twice.