
Querying Infrastructure Monitoring Data With Natural Language
Natural language can work as a front end for infrastructure monitoring and operations data. The system converts an operator's question into structured queries against time-series, consumption, inventory, alarm, and relationship data, then returns the result in a readable form along with where the data came from, which metric was used, what time range was scanned, and how the result was calculated.
The operator asks the operational question directly instead of first working out which dashboard, metric, filter, or report to open. The monitoring system still provides the data; natural language only changes how the user gets to it.
What does a natural-language operations query actually do?
It translates an operator's request into one or more machine-readable requests against operational data. The source operations-assistant design gives a concrete example: "Which nodes exceeded 90% utilization in the last 24 hours and what tasks were running on them?"
That one sentence holds several pieces of query intent. The user wants nodes, the metric is utilization, the threshold is 90%, the time window is the last 24 hours, and the answer also needs workload or task relationships. A useful assistant has to pick out those elements, query the monitoring and consumption data, join the relevant relationships, and return a structured result. The monitoring platform stays in place; the natural-language layer turns an operations question into a monitoring request for it.
What kinds of infrastructure data can be queried this way?
The source design supports natural-language access to monitoring, consumption, configuration, alarm, and operations knowledge, which opens up several kinds of query:
- Resource state: Which nodes are unhealthy? Which accelerator cards are degraded? Which racks are near their power threshold?
- Utilization: Which nodes exceeded a defined utilization level? Which accelerator resources have been idle for a specified period? Which projects are consuming the most compute?
- Operations: Which alarms are still open? Which work orders relate to this device? What happened during the previous incident on this server?
- Capacity: When will the current resource pool reach a defined allocation threshold? Where is resource fragmentation preventing jobs from starting?
- Knowledge: What does this architecture component do? Which runbook applies to this type of incident?
What makes this worth building is that all of these different operational data types sit behind one conversational entry point.
Why is this better than asking users to know metric names?
Operators usually think in operational questions, not in internal metric identifiers. A duty engineer may know that a workload is slow without knowing the exact name of the storage-throughput metric. A manager may want to know when capacity will get tight without knowing which resource-allocation table drives the forecast.
A natural-language layer can translate those questions into the correct operational queries, so not every user has to memorize metric names, dashboard locations, filter syntax, object identifiers, and report definitions.
The source design treats the assistant as the place for asking, while the operations cockpit remains the place for looking. Both interfaces use the same underlying data.
How should the assistant translate a question into data requests?
The translation should identify the object, metric, condition, time range, grouping, and relationship the user asked for. Take this question: "Which GPU nodes had utilization above 90% yesterday and what projects were using them?"
The assistant needs to resolve:
- Object: GPU nodes
- Metric: utilization
- Condition: above 90%
- Time: yesterday
- Relationship: node to task or workload, then workload to project
- Output: qualifying nodes and their project context
A different question, such as "Which services could move from full-card allocation to shared resources?", needs a different data path. Here the assistant needs resource-allocation history, utilization behavior, resource specification, and the policy used to spot possible fragmentation or over-allocation. The source model explicitly lists fragmentation identification and optimization recommendations as assistant capabilities.
Whatever the question, the conversion has to be accurate enough that the final answer matches what the user actually meant.
Why does CMDB relationship data matter?
Natural-language operations questions often need relationships as well as metrics. A monitoring database can tell you that Node A had high utilization, but the user may also ask which workload was running, which project owns it, or which business service depends on it. Those answers come from the relationship data.
The source platform connects tasks to containers, containers to GPUs, GPUs to physical nodes, nodes to network and storage, and services to projects and owners. That lets one natural-language question move across several operational layers.
For the relationship model itself, how a CMDB can connect servers, GPUs, containers, applications, business services, and owners explains how those links support operations.
How should time ranges be handled?
The assistant should resolve the requested time period explicitly and use the same period across related data sources. The period might be the last 15 minutes, the last 24 hours, yesterday, this week, or the period of Incident 1842. If the user asks about utilization and workload assignment over the same 24-hour window, both datasets need to be evaluated the same way.
The source assistant design also puts weight on stating the statistical definition, because "utilization above 90%" can mean several things. Maximum utilization may have crossed 90%, average utilization may have stayed above 90%, or a 95th percentile may have exceeded 90%. The answer should say which interpretation it used, and the statistical meaning should never be hidden behind natural language.
What source attribution should be shown?
Every operational answer should show its source clearly enough to verify. The source design explicitly calls for data-source attribution, metric names, statistical definitions, scan scope, and data granularity. A useful answer might say:
- Data source: node utilization telemetry
- Metric: accelerator utilization
- Time range: previous 24 hours
- Aggregation: five-minute average
- Scan scope: 62 online nodes
- Relationship source: task and project bindings
The display can be compact. What matters is that the operator can understand how the answer was produced, and that matters most for forecasts and recommendations. If an assistant says capacity will reach a threshold in 23 days, the user should be able to see the data and assumptions behind that estimate.
How can natural language be used for incident troubleshooting?
The assistant can combine live operational data with historical operational knowledge. The source knowledge base includes work orders, alarm events, architecture documents, and emergency runbooks, which are ingested and vectorized together.
So a user can ask, "Why is this node unhealthy, and have we seen this before?" The assistant uses the active monitoring data for the current condition and retrieves previous work orders for similar incidents. It can then show the current alarm, a likely cause or investigation path, the historical incident, the relevant runbook, and a recommended next check.
For the broader assistant design, how AI assistants use alarms, work orders, runbooks, and infrastructure documents explains how live queries and knowledge retrieval work together.
How can natural language support capacity forecasting?
The source assistant includes capacity forecasting and expiry-risk reminders. A user can ask, "When will this resource pool reach 90% allocation?" The assistant retrieves the historical allocation trend and applies the platform's forecasting logic.
The result should stay traceable, showing the current allocation, the trend window, the threshold, the forecast date or interval, and the data source.
The same interface can answer questions about fragmentation. Asked "Why are jobs waiting when GPUs are free?", the assistant can inspect the requested resource shape, available devices, topology, and scheduling state, then explain whether the cause is quota, fragmentation, health exclusion, or another known condition.
How can natural language query business consumption?
The same approach works on metering data, with questions like these:
- Which project consumed the most accelerator card hours this month?
- Which model produced the most Tokens today?
- Which tenant has the highest idle allocation?
- Which accelerator type has the highest unit cost?
The source platform allocates consumption across project, tenant, model, and accelerator type, which gives the assistant stable dimensions to query. A user does not need to open the cost dashboard first; they can ask the question and get the grouped result. That result should still use the same definitions as the formal metering dashboard, because natural language should not create a second version of the numbers.
How should the assistant handle ambiguous questions?
When the operational model provides enough context, the assistant should resolve the ambiguity from it and make the chosen interpretation visible.
Take "Show me overloaded nodes." The platform needs a definition of overloaded. If the operations policy defines it as sustained utilization above a specific threshold, the assistant can use that rule and say so. The source materials do not specify a universal threshold, so if no definition exists, the assistant should not invent one. It can ask for a threshold, or present the available utilization data without labeling anything overloaded.
The same goes for terms such as "expensive," "unhealthy," or "high risk." Operational language should map to defined platform rules where they exist.
What should happen when the requested data is unavailable?
The assistant should say it cannot retrieve the requested data. The source assistant design makes this a hard rule: missing data must not be filled with a plausible answer.
A query can fail for several reasons. The metric may not be collected, the requested time range may be outside retention, the user may lack permission, the relationship may be missing, or the data source may be down. The assistant should tell those cases apart when the platform can identify them. An explicit "data unavailable" response is operationally safer than a confident estimate with no evidence.
How should permissions apply to natural-language queries?
Natural-language access should use the same organization, role, tenant, project, and least-privilege controls as the rest of the platform. The source governance design includes a department and team hierarchy, a role permission matrix, tenant membership, project membership, least privilege, and separate authorization for sensitive operations.
The assistant should inherit all of those controls. A user who cannot open Project B's monitoring data through the normal interface should not be able to retrieve it by asking a conversational question. The natural-language layer is one more interface to the same operational data, and it must not become a way around permissions.
Can a natural-language query trigger a change?
The source design keeps a clear execution boundary. The assistant can analyze and recommend, but change actions still need workflow approval before they run.
A user can ask, "Which degraded nodes should be removed from the pool?" and the assistant can return the candidates and the evidence. If the user then wants to make the change, the action goes into the normal authorization process.
For the governance model, how enterprises automate data center operations while keeping approvals, permissions, rollback, and audit controls explains how recommendation and execution stay connected without removing control.
What makes a natural-language operations interface trustworthy?
Three things matter most: the answer has to come from the real operations data, the reasoning path has to be visible enough to verify, and the assistant has to stay inside its permissions and execution boundary.
A platform example that uses natural-language-to-monitoring query conversion with source attribution is Sensaka.
If I were evaluating this capability, I would test questions that require several data sources at once. Ask which nodes exceeded a utilization threshold, what workloads were running, which projects owned them, and whether any related incidents existed. If the assistant returns the correct result with explicit data sources and time scope, natural-language querying is doing more than paraphrasing dashboards.
Frequently Asked Questions
What can an IT operations team ask in natural language?
Teams can ask about resource utilization, abnormal nodes, active workloads, capacity trends, fragmentation, alarms, architecture, and previous incidents when those data sources are connected to the assistant.
How should an AI assistant answer a monitoring-data question?
It should convert the question into a query against the relevant operational data, return a structured result, and show the data source, metric definition, time range, scan scope, and data granularity.
What should happen if the requested operations data cannot be retrieved?
The assistant should say that the data is unavailable instead of inventing an answer. Operational recommendations should remain traceable to retrieved data, documents, alarms, or work orders.