
What KPIs should infrastructure operations teams track in a unified operations dashboard?
Infrastructure operations teams should track KPIs that explain supply, efficiency, service quality, cost, and consumption together. The source operations model uses exactly those five categories because no single metric can explain whether infrastructure is healthy, productive, economical, and delivering the expected service.
A good dashboard should therefore answer more than "Is the hardware up?" It should also answer "How much usable capacity remains, how well is it being used, what quality is being delivered, what does it cost, and who is consuming it?"
What are the five KPI groups in the source model?
The unified operations cockpit groups indicators into five categories:
Resource supply
Production efficiency
Service quality
Operating cost
Business consumption
The internal-share material also uses a four-category summary that combines the operating focus into resource supply, production efficiency, service quality, and operating cost.
The more detailed cockpit adds business consumption as a separate category.
Both views are consistent.
The difference is how much of the consumption side is shown explicitly.
The five-category model is useful because it covers the full operating chain from infrastructure capacity to service output.
What should be measured under resource supply?
Resource supply should show how much capacity exists and how much is usable.
The source examples include:
Total accelerator cards
Available accelerator cards
Total nodes
Online nodes
Cluster capacity distribution
Capacity-level alerts
Forecast expansion date
For broader infrastructure, the same category can include:
Servers
Virtual machines
Storage capacity
Network capacity
Rack space
Power headroom
Cooling headroom
The exact object depends on the environment.
The important distinction is between installed and usable.
A node that is offline, degraded, under maintenance, or blocked by another constraint should not be counted as fully available production capacity.
Why is available capacity more important than total capacity?
Because total asset count can overstate what the business can actually use.
The source AI infrastructure model repeatedly shows that deployable capacity depends on several constraints.
A rack can have U space but lack power.
A GPU can be present but degraded.
A resource pool can have free cards but the wrong topology or resource shape.
A storage pool can have free terabytes but insufficient throughput.
The KPI should therefore distinguish nominal supply from usable supply.
That makes capacity decisions more realistic.
For the physical side, how data centers can detect stranded capacity caused by power, cooling, network, or storage constraints explains why one free-resource number is often misleading.
What should be measured under production efficiency?
Production efficiency shows whether expensive capacity is doing useful work.
The source dashboard includes:
GPU utilization
Running tasks
Queued tasks
Completed tasks
Idle rate
Task success rate
Queue time
Idle reason and recommendations
Those indicators work together.
Utilization alone can be misleading.
A node can show high utilization while tasks fail repeatedly.
A cluster can show low utilization because there is little demand.
It can also show low utilization because storage is slow or resources are fragmented.
That is why the source model pairs utilization with queue and task metrics.
Efficiency should be interpreted in context.
Why should idle rate be a KPI?
Idle rate converts unused capacity into an operating signal.
The source cockpit includes idle rate in the cost section.
The detailed operations model also analyzes idle reasons and provides recommendations.
This is important because idle capacity can have several causes.
No demand.
Oversized allocation.
Storage bottleneck.
Network bottleneck.
Resource fragmentation.
Health exclusion.
Quota.
The KPI should therefore lead to the cause rather than stop at the percentage.
A high idle rate is a symptom.
The action depends on why the resources are idle.
What should be measured under service quality?
Service quality should measure whether the delivered service is meeting its expected operating standard.
The source examples include:
Token success rate
Task success rate
Pending alarms
Healthy, warning, and failed nodes
SLO
Error budget
Failure detection time
Recovery time
This category connects infrastructure operations to service delivery.
A resource can be online while service quality is poor.
For example:
All nodes online.
Token success rate falling.
Latency rising.
The device-availability KPI looks healthy.
The service is not.
That is why service-level indicators belong next to infrastructure indicators.
Why should MTTD and MTTR be included?
Because availability alone does not show how effectively the operations team handles incidents.
The source SRE framework includes failure detection and recovery duration.
MTTD measures how quickly a problem becomes visible.
MTTR measures how quickly the service is restored under the organization's definition.
Those indicators show whether monitoring, diagnosis, workflow, and automation are improving.
If incident count stays similar but MTTR falls significantly, operations are becoming more effective.
For the full framework, how SRE metrics such as MTTD, MTTR, SLO, error budgets, and burn rate apply to AI and infrastructure operations explains how the reliability metrics fit together.
What should be measured under operating cost?
Operating cost should connect capacity consumption to financial impact.
The source cockpit includes:
Monthly accelerator card hours
Idle rate
Unit Token cost
The internal-share version also tracks:
Cost per thousand Tokens
Project billing
Tenant billing
The broader metering model connects accelerator hours, energy, Tokens, projects, tenants, and models.
That gives the dashboard several useful cost views.
Total consumption.
Unit cost.
Idle cost.
Cost by owner.
The important requirement is consistent definitions.
If card hours mean allocated time in one page and active time in another, the KPI becomes unreliable.
What should be measured under business consumption?
Business consumption shows how the infrastructure is being used by projects, tenants, models, or applications.
The source cockpit includes Token volume and quota views.
The metering layer can also show:
Project usage
Tenant usage
Model usage
API calls
Accelerator card hours
This category is important because operations does not exist only to keep infrastructure healthy.
It exists to deliver capacity to users and services.
Consumption KPIs show whether demand is changing and which parts of the business are driving it.
They also support planning and chargeback.
What capacity-risk KPIs should be included?
The source operations model includes:
Capacity-level warning
Forecast expansion date
Capacity prediction
Fragmentation identification
Expiry risk
That means the dashboard should include forward-looking indicators, not only current utilization.
Examples include:
Current allocation percentage
Available capacity
Threshold warning
Forecast threshold date
Resource fragmentation
Physical bottleneck
Contract or support expiry
The exact forecast method is not defined in the source.
So the KPI should not claim precision beyond the underlying model.
What matters is that the dashboard helps operators see approaching constraints early enough to act.
What alarm KPIs should be included?
The source cockpit shows:
Pending alarms
Latest alarms
Node-state exceptions
The AIOps layer adds incident consolidation and root-cause analysis.
The top-level dashboard does not need every raw alarm count.
It needs enough information to show current operational risk.
Useful indicators include:
Open incidents
Pending critical alarms
Affected nodes or services
Unresolved duration
Repeated incidents
Raw event volume can still be available in the alarm center.
The dashboard should focus attention on what requires action.
Should infrastructure KPIs include business-service impact?
Yes, when the relationship data is available.
The source data foundation supports business topology and impact analysis.
That means a KPI can distinguish between:
Four failed devices with no active service impact.
One failed device supporting a critical service.
The second situation may deserve higher priority.
Infrastructure operations become more useful when device state is connected to business-service health.
For the service layer, what is business service management, and how is it different from infrastructure monitoring explains why impact and ownership matter.
How should energy KPIs fit into the dashboard?
The source infrastructure model includes:
Power consumption
PUE
WUE
Unit Token energy
Peak and off-peak pricing
Energy cost
These indicators belong in the operating-cost and facility-efficiency views.
They can also support capacity planning.
High rack power can reduce deployable capacity even when U space remains.
Liquid-cooling and power constraints can therefore appear both as facility KPIs and as capacity-risk indicators.
The dashboard should preserve that relationship instead of showing energy as an isolated sustainability page.
How should KPI definitions be governed?
Use one shared metric dictionary.
The source cockpit says its figures come from the existing collection and metering systems and are not manually re-entered.
That is the right rule.
For every KPI, define:
Name
Source
Calculation
Time window
Aggregation method
Unit
Owner
Drill-down target
This prevents different teams from creating several versions of "utilization" or "available capacity."
The cockpit should show the agreed operational definition.
How many KPIs should the first screen show?
The source design favors a small number of categories and then drill-down.
That is a useful principle.
The first screen should show enough to understand current operations quickly.
Detailed component metrics belong in specialist views.
A center leader does not need every storage counter on the first screen.
A storage engineer still needs those counters after drill-down.
The dashboard should therefore prioritize the indicators that change decisions.
More metrics do not automatically create more visibility.
How should KPIs differ by role?
The source model identifies different priorities by user.
Center leader:
Total supply
Trends
Capacity risk
Cost
Operations role:
Efficiency
Idle rate
Cost
Capacity
Duty role:
Node state
Current alarms
Failures
Recovery
Business user:
Consumption
Quota
Service quality
The underlying definitions stay the same.
The presentation changes by role.
That reduces clutter without creating separate versions of the truth.
What is the minimum useful KPI set?
A source-grounded minimum could include:
Available capacity
Online node count
Utilization
Running and queued tasks
Task or service success rate
Open operational alarms
Idle rate
Unit cost
Consumption volume
Capacity threshold date
Those indicators cover supply, efficiency, reliability, cost, consumption, and future risk.
A platform example that organizes operations around these KPI categories is Sensaka.
If I were designing the first version of a unified operations dashboard, I would start with one metric from each decision category rather than 50 domain metrics. Show usable supply, utilization, service success, current operational risk, unit cost, and business consumption. Then make every KPI drill down to the evidence behind it.
Frequently Asked Questions
What are the main KPI groups in the source operations model?
The source cockpit groups KPIs into resource supply, production efficiency, service quality, operating cost, and business consumption.
Should a dashboard show only utilization and uptime?
No. The source model also tracks queue state, success rate, idle rate, unit cost, Token consumption, alarms, SLOs, and capacity risk so operations can measure productivity and service delivery, not only device availability.
How should KPI definitions be governed?
The source design says cockpit KPIs should use the same definitions and data sources as their detailed pages so the organization does not create conflicting numbers.