
How AI Assistants Can Use Alarms, Work Orders and Runbooks in IT Ops
AI assistants can help IT operations teams by combining historical knowledge and live operational data. Alarms explain what is happening now, work orders show what happened before, runbooks describe approved response procedures, architecture documents provide system context, and monitoring data verifies the current state.
The source operations-assistant design draws a firm boundary: the assistant can query, troubleshoot, predict, and recommend, but infrastructure changes still go through authorization workflows. That keeps the assistant useful without turning natural-language interaction into an uncontrolled execution path.
What is an AI operations assistant?
An AI operations assistant is a conversational interface over operational knowledge and data. It differs from a generic chatbot because the useful answers depend on the organization's actual infrastructure. A generic model may know what an ECC error is. An operations assistant has to say which card has the ECC error, which server contains it, and which workload is using it. It also has to know whether this happened before, which runbook applies, what happened the last time, whether a work order is already open, and whether the workload can move to another resource.
The source model calls this the difference between a cockpit and an assistant: the cockpit is for looking, and the assistant is for asking.
What information should go into the operations knowledge base?
The source assistant design explicitly combines historical work orders, alarm events, architecture documents, and emergency runbooks. Those sources cover different kinds of knowledge. Work orders contain experience, alarms contain event history, architecture documents contain structure and design intent, and runbooks contain approved response steps.
The source design also describes the content as ingested and vectorized together so it can be retrieved through natural-language questions. The assistant can then search across these materials, and the operator no longer has to remember which system contains the answer.
Why are work orders useful to an AI assistant?
Work orders contain the organization's real operational history. A work order can show what failed, who handled it, what diagnosis was made, what action was taken, which spare part was used, how long recovery took, whether the fix worked, and what post-incident note was added. That makes work orders useful for troubleshooting.
Suppose the current alarm resembles an incident from two months ago. The assistant can retrieve the previous work order and explain the similarity, but it should not blindly assume the same root cause. The current evidence still needs to be checked, because historical work is guidance and does not prove anything about today's failure.
The source model also says the knowledge base continues learning from new work orders, so the operational memory improves over time.
Why are alarms useful to the assistant?
Alarms provide the current event context. A user may ask why Node 17 is unhealthy, and the assistant can retrieve the active alarms for that node. If the platform consolidates raw alarms into an incident, the assistant can use the incident instead of reading 37 unrelated messages one by one.
The source AIOps design includes alarm aggregation, downstream suppression, root-cause evidence, and incident timelines. That gives the assistant a cleaner event object, so it can explain what triggered first, which alarms are likely downstream symptoms, what component is suspected, and what confidence or evidence supports the conclusion.
For the event-correlation layer, how AIOps reduces alarm noise, identifies root causes, and determines business impact explains the full flow.
Why are runbooks useful?
Runbooks describe the organization's approved response to known situations: restarting a failed collector, quarantining a degraded GPU, recovering a training job from checkpoint, replacing a failed power supply, escalating a private-line outage, or handling a liquid-cooling warning.
The source incident workflow explicitly includes runbook matching and recommended actions. The assistant can therefore tell the operator which runbook matches this incident, what the required steps are, and which steps are automatic, which require approval, and which require a human.
The runbook keeps the answer consistent with the organization's operating procedure, so the assistant does not generate a new procedure from scratch every time.
Why are architecture documents useful?
Architecture documents provide context that raw monitoring cannot always infer. A document can explain why two services are connected, which network path is primary and which is backup, what one cluster is intended to do, which failure domains matter, how capacity is designed, and which system owns a function.
The source assistant design explicitly includes architecture documents in the operations knowledge base. This matters because topology discovery can show that two systems communicate, but it may not explain the design intent. Architecture documentation can.
The assistant should retrieve the relevant part of the document and keep it apart from live state. "Architecture says this is the backup path" is different from "monitoring says the backup path is currently active."
How can the assistant answer live-data questions?
The source design converts natural-language requests into monitoring-data queries. That lets the operator ask a question such as:
"Which nodes exceeded 90 percent utilization in the last 24 hours and what tasks were running on them?"
The assistant turns the request into structured queries against time-series and consumption data and returns structured results.
The control that matters most here is source attribution. The source material says answers should include the data source, metric name, statistical definition, scan scope, and data granularity where relevant. That makes the answer verifiable. The assistant should not invent a utilization number because it sounds plausible.
What should source attribution look like?
Source attribution should tell the operator where the answer came from. Useful details include the metric source and name, the time range, the aggregation method, scan scope, data granularity, and whichever knowledge document, work order, alarm, or incident the answer used.
The source assistant example includes metric-level attribution, 15-second granularity, and scan scope, which is a strong trust design. If the operator asks why the assistant says allocation will reach 90 percent in 23 days, they should be able to see which capacity data and forecast window produced the estimate. An answer that cannot be traced back to operational evidence should not be treated as an operations decision.
What should the assistant do when data is unavailable?
It should say that the data cannot be retrieved. The source design makes this an explicit rule, and the assistant should not fill a monitoring gap with a guess.
If GPU power telemetry is unavailable, it should say it cannot calculate energy from that source. If a work order does not exist, it should not invent a previous repair. If the architecture document does not describe the requested dependency, it should say the documentation does not support the conclusion.
This behavior matters even more in operations, because a confident invented answer can cause a real change to be made on the wrong resource.
How can the assistant help troubleshoot incidents?
A useful troubleshooting flow can combine five kinds of evidence: the current alarm, the affected resource, related topology, historical work orders, and the matching runbook. From those, the assistant can produce a structured explanation covering the observed problem, the likely cause, the evidence, affected services, historical similarity, the recommended runbook, and the next verification step.
The source incident model also provides root-cause confidence and evidence item by item, and the assistant can surface that evidence conversationally. For example:
"The likely root cause is GPU card 3 because ECC errors increased before the worker restart, all affected containers were bound to the same card, and peer cards remained healthy."
That is more useful than saying "GPU issue detected."
How can the assistant use work orders for recommendations?
Historical work orders can reveal which actions succeeded in similar cases. Suppose three previous incidents with the same switch model and alarm sequence were resolved by replacing an optical module. The assistant can mention that pattern, but it should also compare the current evidence, since a different port or a different error pattern may require another action. The recommendation should be framed as evidence-based guidance and should not claim automatic certainty.
The source operations model keeps post-incident reviews in the knowledge base, which gives the assistant better material than raw work-order closure text. A review can explain why the incident occurred and what should be improved.
How can the assistant help with capacity planning?
The source assistant design includes capacity forecasting and fragmentation recommendations. It can answer when the current compute pool will reach the allocation threshold, which resource types are frequently idle, where capacity is stranded, which workloads could move from full-card allocation to shared resources, and which contracts are approaching expiry.
The assistant can turn a management question into a data query and then explain the result. The forecast has to stay connected to the source data and assumptions, so it should not show up as a mysterious AI prediction.
How can the assistant identify resource fragmentation?
Resource fragmentation occurs when total free capacity exists but cannot satisfy the requested resource shape. The assistant can combine resource pool data, queue reasons, GPU models, topology, quota, and workload requests to answer a question like "Why is Job A waiting when 12 GPUs are free?"
The response may explain that the job requires eight matching GPUs in one topology domain while the free cards are distributed across incompatible nodes. That kind of conversational diagnosis helps because it turns scheduler state into an explanation an operator can act on.
Can an AI assistant create a work order?
It can create a proposed work order, or convert a recommendation into the normal workflow when the platform permits it. The source assistant example says recommendations can be turned into work orders.
That is a sensible boundary. The assistant helps move from analysis to process, while the workflow remains responsible for approval and execution. The audit trail is preserved, and the operator does not need to retype the incident into another system.
Should an AI assistant execute changes directly?
The safer default in the source design is no. The assistant provides analysis and recommended actions, and change actions still require workflow approval before execution. The source material repeats this boundary several times, deliberately.
A natural-language interface should not bypass permissions, risk classification, approval, canary rollout, rollback, or audit.
For the control model, how enterprises automate data center operations while keeping approvals, permissions, rollback, and audit controls explains how recommendation and execution should remain connected but governed.
How should the assistant handle permissions?
The assistant should respect the same organizational and tenant boundaries as the underlying platform. A user should not gain access to restricted operational data simply because they ask for it conversationally.
The source governance model includes departments, roles, tenant membership, project membership, and least-privilege permissions, and those controls should apply to assistant retrieval too. If the user can view only Project A, the assistant should not retrieve Project B's work orders or infrastructure documents. The AI layer should inherit permissions instead of creating a parallel access model.
What should an AI operations assistant not do?
It should not invent live metrics or work-order history, hide uncertainty, ignore source boundaries, or bypass approvals. It should not execute high-risk actions from plain text alone, return restricted data, or present a recommendation as confirmed root cause without evidence.
The source assistant design also includes clear refusal of out-of-scope requests. I like that, because an assistant becomes more trustworthy when it knows when not to answer.
What should a useful assistant interface show?
The conversation should remain connected to evidence. A useful response can include a direct answer, a structured table, a mini trend chart, source attribution, related alarms and work orders, the matched runbook, a recommended action, a forecast, and a button or action to create a work order.
The source interface combines natural-language query results, source attribution, and predictive recommendations. That design works because the answer stays verifiable and actionable.
A platform example that uses work orders, alarms, architecture documents, runbooks, and live monitoring data in an AI operations assistant is Sensaka.
If I were deploying an AI assistant for IT operations, I would judge it on one test: ask a question whose answer exists in live monitoring and historical work orders, then verify that the assistant returns the correct current data, cites its operational source, finds the relevant past incident, recommends the approved runbook, and refuses to execute the change without authorization. That is worth far more than a chatbot that can merely explain IT terminology.
Frequently Asked Questions
What data should an AI operations assistant use?
Useful sources include alarms, historical work orders, architecture documents, emergency runbooks, CMDB relationships, monitoring metrics, capacity data, and operational policies.
Should an AI operations assistant execute infrastructure changes directly?
The safer default is analysis and recommendation first, with infrastructure changes still passing through the normal authorization and workflow controls.
How can an AI assistant avoid inventing operational data?
Operational answers should query trusted live data sources, show source attribution and scan scope, and explicitly say when the requested data cannot be retrieved.