Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    KPI
    Infrastructure Operations
    Operations Dashboard

    What KPIs should infrastructure operations teams track in a unified operations dashboard?

    June 15, 2026
    10 min read read

    Infrastructure operations teams should track KPIs that explain supply, efficiency, service quality, cost, and consumption together. The source operations model uses exactly those five categories because no single metric can explain whether infrastructure is healthy, productive, economical, and delivering the expected service.

    A good dashboard should therefore answer more than "Is the hardware up?" It should also answer "How much usable capacity remains, how well is it being used, what quality is being delivered, what does it cost, and who is consuming it?"

    What are the five KPI groups in the source model?

    The unified operations cockpit groups indicators into five categories:

    Resource supply
    Production efficiency
    Service quality
    Operating cost
    Business consumption

    The internal-share material also uses a four-category summary that combines the operating focus into resource supply, production efficiency, service quality, and operating cost.

    The more detailed cockpit adds business consumption as a separate category.

    Both views are consistent.

    The difference is how much of the consumption side is shown explicitly.

    The five-category model is useful because it covers the full operating chain from infrastructure capacity to service output.

    What should be measured under resource supply?

    Resource supply should show how much capacity exists and how much is usable.

    The source examples include:

    Total accelerator cards
    Available accelerator cards
    Total nodes
    Online nodes
    Cluster capacity distribution
    Capacity-level alerts
    Forecast expansion date

    For broader infrastructure, the same category can include:

    Servers
    Virtual machines
    Storage capacity
    Network capacity
    Rack space
    Power headroom
    Cooling headroom

    The exact object depends on the environment.

    The important distinction is between installed and usable.

    A node that is offline, degraded, under maintenance, or blocked by another constraint should not be counted as fully available production capacity.

    Why is available capacity more important than total capacity?

    Because total asset count can overstate what the business can actually use.

    The source AI infrastructure model repeatedly shows that deployable capacity depends on several constraints.

    A rack can have U space but lack power.

    A GPU can be present but degraded.

    A resource pool can have free cards but the wrong topology or resource shape.

    A storage pool can have free terabytes but insufficient throughput.

    The KPI should therefore distinguish nominal supply from usable supply.

    That makes capacity decisions more realistic.

    For the physical side, how data centers can detect stranded capacity caused by power, cooling, network, or storage constraints explains why one free-resource number is often misleading.

    What should be measured under production efficiency?

    Production efficiency shows whether expensive capacity is doing useful work.

    The source dashboard includes:

    GPU utilization
    Running tasks
    Queued tasks
    Completed tasks
    Idle rate
    Task success rate
    Queue time
    Idle reason and recommendations

    Those indicators work together.

    Utilization alone can be misleading.

    A node can show high utilization while tasks fail repeatedly.

    A cluster can show low utilization because there is little demand.

    It can also show low utilization because storage is slow or resources are fragmented.

    That is why the source model pairs utilization with queue and task metrics.

    Efficiency should be interpreted in context.

    Why should idle rate be a KPI?

    Idle rate converts unused capacity into an operating signal.

    The source cockpit includes idle rate in the cost section.

    The detailed operations model also analyzes idle reasons and provides recommendations.

    This is important because idle capacity can have several causes.

    No demand.

    Oversized allocation.

    Storage bottleneck.

    Network bottleneck.

    Resource fragmentation.

    Health exclusion.

    Quota.

    The KPI should therefore lead to the cause rather than stop at the percentage.

    A high idle rate is a symptom.

    The action depends on why the resources are idle.

    What should be measured under service quality?

    Service quality should measure whether the delivered service is meeting its expected operating standard.

    The source examples include:

    Token success rate
    Task success rate
    Pending alarms
    Healthy, warning, and failed nodes
    SLO
    Error budget
    Failure detection time
    Recovery time

    This category connects infrastructure operations to service delivery.

    A resource can be online while service quality is poor.

    For example:

    All nodes online.

    Token success rate falling.

    Latency rising.

    The device-availability KPI looks healthy.

    The service is not.

    That is why service-level indicators belong next to infrastructure indicators.

    Why should MTTD and MTTR be included?

    Because availability alone does not show how effectively the operations team handles incidents.

    The source SRE framework includes failure detection and recovery duration.

    MTTD measures how quickly a problem becomes visible.

    MTTR measures how quickly the service is restored under the organization's definition.

    Those indicators show whether monitoring, diagnosis, workflow, and automation are improving.

    If incident count stays similar but MTTR falls significantly, operations are becoming more effective.

    For the full framework, how SRE metrics such as MTTD, MTTR, SLO, error budgets, and burn rate apply to AI and infrastructure operations explains how the reliability metrics fit together.

    What should be measured under operating cost?

    Operating cost should connect capacity consumption to financial impact.

    The source cockpit includes:

    Monthly accelerator card hours
    Idle rate
    Unit Token cost

    The internal-share version also tracks:

    Cost per thousand Tokens
    Project billing
    Tenant billing

    The broader metering model connects accelerator hours, energy, Tokens, projects, tenants, and models.

    That gives the dashboard several useful cost views.

    Total consumption.

    Unit cost.

    Idle cost.

    Cost by owner.

    The important requirement is consistent definitions.

    If card hours mean allocated time in one page and active time in another, the KPI becomes unreliable.

    What should be measured under business consumption?

    Business consumption shows how the infrastructure is being used by projects, tenants, models, or applications.

    The source cockpit includes Token volume and quota views.

    The metering layer can also show:

    Project usage
    Tenant usage
    Model usage
    API calls
    Accelerator card hours

    This category is important because operations does not exist only to keep infrastructure healthy.

    It exists to deliver capacity to users and services.

    Consumption KPIs show whether demand is changing and which parts of the business are driving it.

    They also support planning and chargeback.

    What capacity-risk KPIs should be included?

    The source operations model includes:

    Capacity-level warning
    Forecast expansion date
    Capacity prediction
    Fragmentation identification
    Expiry risk

    That means the dashboard should include forward-looking indicators, not only current utilization.

    Examples include:

    Current allocation percentage
    Available capacity
    Threshold warning
    Forecast threshold date
    Resource fragmentation
    Physical bottleneck
    Contract or support expiry

    The exact forecast method is not defined in the source.

    So the KPI should not claim precision beyond the underlying model.

    What matters is that the dashboard helps operators see approaching constraints early enough to act.

    What alarm KPIs should be included?

    The source cockpit shows:

    Pending alarms
    Latest alarms
    Node-state exceptions

    The AIOps layer adds incident consolidation and root-cause analysis.

    The top-level dashboard does not need every raw alarm count.

    It needs enough information to show current operational risk.

    Useful indicators include:

    Open incidents
    Pending critical alarms
    Affected nodes or services
    Unresolved duration
    Repeated incidents

    Raw event volume can still be available in the alarm center.

    The dashboard should focus attention on what requires action.

    Should infrastructure KPIs include business-service impact?

    Yes, when the relationship data is available.

    The source data foundation supports business topology and impact analysis.

    That means a KPI can distinguish between:

    Four failed devices with no active service impact.

    One failed device supporting a critical service.

    The second situation may deserve higher priority.

    Infrastructure operations become more useful when device state is connected to business-service health.

    For the service layer, what is business service management, and how is it different from infrastructure monitoring explains why impact and ownership matter.

    How should energy KPIs fit into the dashboard?

    The source infrastructure model includes:

    Power consumption
    PUE
    WUE
    Unit Token energy
    Peak and off-peak pricing
    Energy cost

    These indicators belong in the operating-cost and facility-efficiency views.

    They can also support capacity planning.

    High rack power can reduce deployable capacity even when U space remains.

    Liquid-cooling and power constraints can therefore appear both as facility KPIs and as capacity-risk indicators.

    The dashboard should preserve that relationship instead of showing energy as an isolated sustainability page.

    How should KPI definitions be governed?

    Use one shared metric dictionary.

    The source cockpit says its figures come from the existing collection and metering systems and are not manually re-entered.

    That is the right rule.

    For every KPI, define:

    Name
    Source
    Calculation
    Time window
    Aggregation method
    Unit
    Owner
    Drill-down target

    This prevents different teams from creating several versions of "utilization" or "available capacity."

    The cockpit should show the agreed operational definition.

    How many KPIs should the first screen show?

    The source design favors a small number of categories and then drill-down.

    That is a useful principle.

    The first screen should show enough to understand current operations quickly.

    Detailed component metrics belong in specialist views.

    A center leader does not need every storage counter on the first screen.

    A storage engineer still needs those counters after drill-down.

    The dashboard should therefore prioritize the indicators that change decisions.

    More metrics do not automatically create more visibility.

    How should KPIs differ by role?

    The source model identifies different priorities by user.

    Center leader:

    Total supply
    Trends
    Capacity risk
    Cost

    Operations role:

    Efficiency
    Idle rate
    Cost
    Capacity

    Duty role:

    Node state
    Current alarms
    Failures
    Recovery

    Business user:

    Consumption
    Quota
    Service quality

    The underlying definitions stay the same.

    The presentation changes by role.

    That reduces clutter without creating separate versions of the truth.

    What is the minimum useful KPI set?

    A source-grounded minimum could include:

    Available capacity
    Online node count
    Utilization
    Running and queued tasks
    Task or service success rate
    Open operational alarms
    Idle rate
    Unit cost
    Consumption volume
    Capacity threshold date

    Those indicators cover supply, efficiency, reliability, cost, consumption, and future risk.

    A platform example that organizes operations around these KPI categories is Sensaka.

    If I were designing the first version of a unified operations dashboard, I would start with one metric from each decision category rather than 50 domain metrics. Show usable supply, utilization, service success, current operational risk, unit cost, and business consumption. Then make every KPI drill down to the evidence behind it.

    Frequently Asked Questions

    What are the main KPI groups in the source operations model?

    The source cockpit groups KPIs into resource supply, production efficiency, service quality, operating cost, and business consumption.

    Should a dashboard show only utilization and uptime?

    No. The source model also tracks queue state, success rate, idle rate, unit cost, Token consumption, alarms, SLOs, and capacity risk so operations can measure productivity and service delivery, not only device availability.

    How should KPI definitions be governed?

    The source design says cockpit KPIs should use the same definitions and data sources as their detailed pages so the organization does not create conflicting numbers.