
KPIs to Track in a Unified Infrastructure Operations Dashboard
Infrastructure operations teams should track KPIs that explain supply, efficiency, service quality, cost, and consumption together. The source operations model uses exactly those five categories because no single metric can explain whether infrastructure is healthy, productive, economical, and delivering the expected service.
A good dashboard should therefore answer more than "Is the hardware up?" It should also answer "How much usable capacity remains, how well is it being used, what quality is being delivered, what does it cost, and who is consuming it?"
What are the five KPI groups in the source model?
The unified operations cockpit groups indicators into five categories: resource supply, production efficiency, service quality, operating cost, and business consumption.
The internal-share material also uses a four-category summary that combines the operating focus into resource supply, production efficiency, service quality, and operating cost, while the more detailed cockpit adds business consumption as a separate category. Both views are consistent and differ only in how much of the consumption side is shown explicitly. The five-category model is useful because it covers the full operating chain from infrastructure capacity to service output.
What should be measured under resource supply?
Resource supply should show how much capacity exists and how much is usable. The source examples include total and available accelerator cards, total and online nodes, cluster capacity distribution, capacity-level alerts, and the forecast expansion date.
For broader infrastructure, the same category can include servers, virtual machines, storage capacity, network capacity, rack space, power headroom, and cooling headroom. The exact object depends on the environment.
The distinction that matters is between installed and usable. A node that is offline, degraded, under maintenance, or blocked by another constraint should not be counted as fully available production capacity.
Why is available capacity more important than total capacity?
Total asset count can overstate what the business can actually use. The source AI infrastructure model repeatedly shows that deployable capacity depends on several constraints. A rack can have U space but lack power. A GPU can be present but degraded. A resource pool can have free cards with the wrong topology or resource shape, and a storage pool can have free terabytes but insufficient throughput.
The KPI should therefore distinguish nominal supply from usable supply, which makes capacity decisions more realistic.
For the physical side, how data centers can detect stranded capacity caused by power, cooling, network, or storage constraints explains why one free-resource number is often misleading.
What should be measured under production efficiency?
Production efficiency shows whether expensive capacity is doing useful work. The source dashboard includes GPU utilization, running, queued, and completed tasks, idle rate, task success rate, queue time, and idle reasons with recommendations.
Those indicators work together, because utilization alone can mislead. A node can show high utilization while tasks fail repeatedly. A cluster can show low utilization because there is little demand, or because storage is slow or resources are fragmented. That is why the source model pairs utilization with queue and task metrics, and why efficiency should be read in context.
Why should idle rate be a KPI?
Idle rate converts unused capacity into an operating signal. The source cockpit includes idle rate in the cost section, and the detailed operations model also analyzes idle reasons and provides recommendations.
This matters because idle capacity can have several causes: no demand, oversized allocation, a storage or network bottleneck, resource fragmentation, health exclusion, or quota. The KPI should lead to the cause instead of stopping at the percentage. A high idle rate is a symptom, and the action depends on why the resources are idle.
What should be measured under service quality?
Service quality should measure whether the delivered service meets its expected operating standard. The source examples include Token success rate, task success rate, pending alarms, healthy, warning, and failed nodes, SLO, error budget, failure detection time, and recovery time.
This category connects infrastructure operations to service delivery, since a resource can be online while service quality is poor. Picture all nodes online while the Token success rate falls and latency rises. The device-availability KPI looks healthy even though the service is struggling. That is why service-level indicators belong next to infrastructure indicators.
Why should MTTD and MTTR be included?
Availability alone does not show how effectively the operations team handles incidents. The source SRE framework includes failure detection and recovery duration.
MTTD measures how quickly a problem becomes visible, and MTTR measures how quickly the service is restored under the organization's definition. Together they show whether monitoring, diagnosis, workflow, and automation are improving. If incident count stays similar but MTTR falls significantly, operations are becoming more effective.
For the full framework, how SRE metrics such as MTTD, MTTR, SLO, error budgets, and burn rate apply to AI and infrastructure operations explains how the reliability metrics fit together.
What should be measured under operating cost?
Operating cost should connect capacity consumption to financial impact. The source cockpit includes monthly accelerator card hours, idle rate, and unit Token cost. The internal-share version also tracks cost per thousand Tokens, project billing, and tenant billing, and the broader metering model connects accelerator hours, energy, Tokens, projects, tenants, and models.
That gives the dashboard several cost views: total consumption, unit cost, idle cost, and cost by owner. They only work with consistent definitions. If card hours mean allocated time in one page and active time in another, the KPI becomes unreliable.
What should be measured under business consumption?
Business consumption shows how projects, tenants, models, or applications use the infrastructure. The source cockpit includes Token volume and quota views, and the metering layer can also show project, tenant, and model usage, API calls, and accelerator card hours.
This category matters because operations exists to deliver capacity to users and services as well as to keep infrastructure healthy. Consumption KPIs show whether demand is changing and which parts of the business are driving it, and they support planning and chargeback.
What capacity-risk KPIs should be included?
The source operations model includes capacity-level warnings, the forecast expansion date, capacity prediction, fragmentation identification, and expiry risk. The dashboard should therefore include forward-looking indicators as well as current utilization. Examples include the current allocation percentage, available capacity, threshold warnings, the forecast threshold date, resource fragmentation, physical bottlenecks, and contract or support expiry.
The source does not define the exact forecast method, so the KPI should not claim precision beyond the underlying model. What matters is that the dashboard helps operators see approaching constraints early enough to act.
What alarm KPIs should be included?
The source cockpit shows pending alarms, latest alarms, and node-state exceptions, and the AIOps layer adds incident consolidation and root-cause analysis.
The top-level dashboard does not need every raw alarm count. It needs enough information to show current operational risk: open incidents, pending critical alarms, affected nodes or services, unresolved duration, and repeated incidents. Raw event volume can still be available in the alarm center, while the dashboard should focus attention on what requires action.
Should infrastructure KPIs include business-service impact?
Yes, when the relationship data is available. The source data foundation supports business topology and impact analysis, so a KPI can distinguish four failed devices with no active service impact from one failed device supporting a critical service. The second situation may deserve higher priority. Infrastructure operations become more useful when device state is connected to business-service health.
For the service layer, what is business service management, and how is it different from infrastructure monitoring explains why impact and ownership matter.
How should energy KPIs fit into the dashboard?
The source infrastructure model includes power consumption, PUE, WUE, unit Token energy, peak and off-peak pricing, and energy cost. These indicators belong in the operating-cost and facility-efficiency views, and they can also support capacity planning.
High rack power can reduce deployable capacity even when U space remains, so liquid-cooling and power constraints can appear both as facility KPIs and as capacity-risk indicators. The dashboard should preserve that relationship instead of showing energy as an isolated sustainability page.
How should KPI definitions be governed?
Use one shared metric dictionary. The source cockpit says its figures come from the existing collection and metering systems and are not manually re-entered, which is the right rule.
For every KPI, define the name, source, calculation, time window, aggregation method, unit, owner, and drill-down target. This prevents different teams from creating several versions of "utilization" or "available capacity," and the cockpit should show the agreed operational definition.
How many KPIs should the first screen show?
The source design favors a small number of categories and then drill-down, which is a sound principle. The first screen should show enough to understand current operations quickly, and detailed component metrics belong in specialist views. A center leader does not need every storage counter on the first screen, but a storage engineer still needs those counters after drill-down.
The dashboard should prioritize the indicators that change decisions, because more metrics do not automatically create more visibility.
How should KPIs differ by role?
The source model identifies different priorities by user. A center leader cares about total supply, trends, capacity risk, and cost. The operations role watches efficiency, idle rate, cost, and capacity. The duty role follows node state, current alarms, failures, and recovery. A business user looks at consumption, quota, and service quality.
The underlying definitions stay the same while the presentation changes by role, which reduces clutter without creating separate versions of the truth.
What is the minimum useful KPI set?
A source-grounded minimum could include available capacity, online node count, utilization, running and queued tasks, task or service success rate, open operational alarms, idle rate, unit cost, consumption volume, and the capacity threshold date. Those indicators cover supply, efficiency, reliability, cost, consumption, and future risk.
A platform example that organizes operations around these KPI categories is Sensaka.
If I were designing the first version of a unified operations dashboard, I would start with one metric from each decision category instead of 50 domain metrics. Show usable supply, utilization, service success, current operational risk, unit cost, and business consumption. Then make every KPI drill down to the evidence behind it.
Frequently Asked Questions
What are the main KPI groups in the source operations model?
The source cockpit groups KPIs into resource supply, production efficiency, service quality, operating cost, and business consumption.
Should a dashboard show only utilization and uptime?
No. The source model also tracks queue state, success rate, idle rate, unit cost, Token consumption, alarms, SLOs, and capacity risk so operations can measure productivity and service delivery, not only device availability.
How should KPI definitions be governed?
The source design says cockpit KPIs should use the same definitions and data sources as their detailed pages so the organization does not create conflicting numbers.