
Root Cause Analysis With Topology, Metrics, Incidents and CMDB Data
Root cause analysis becomes more reliable when different evidence types answer different parts of the problem. Topology narrows which components could affect the service. Time-series metrics establish what changed first. Historical incidents show whether the pattern has happened before, and configuration relationships explain what depends on what and whether the environment changed around the time of failure.
The source design combines those evidence types into a root-cause conclusion with confidence and itemized reasons. The requirement that matters most is explainability: operators should be able to inspect the evidence instead of receiving an unexplained score.
What question should root cause analysis answer?
Root cause analysis should answer which component or condition most plausibly initiated the incident. It should also say what evidence supports the conclusion, what services are affected, what changed around the incident, and what should be investigated next.
The source closed-loop design does not treat root cause as a single alarm label. It uses relationships, time, historical cases, rules and current evidence. That matters because infrastructure incidents often cross domains. A GPU problem can create container symptoms, a storage problem can look like compute underutilization, and a network problem can look like application failure.
What does topology contribute?
Topology restricts the search to dependencies that could plausibly cause the observed symptoms.
Suppose five application instances fail. The topology shows they run on different containers and several servers, but all those servers depend on one switch, so that switch becomes a strong shared-dependency candidate. In another case, the failed services use different network paths but share one storage system, and storage becomes the stronger candidate.
Topology is useful because it turns a list of alarms into a dependency graph. The source business and infrastructure topology model supports exactly this upward and downward traversal.
What does time-series data contribute?
Time-series data provides sequence and behavior around the incident. Sequence does not prove causality by itself, but it is strong evidence.
The source internal-share example shows ECC count changes before a container restart, and the affected containers share the same card, so the timing supports the hardware hypothesis. Another case can show the opposite pattern: storage latency rises, data-loading wait increases, and GPU utilization falls afterward. That sequence points toward the storage path and away from a GPU defect.
The point is to put cross-domain metrics on the same timeline.
Why should metrics be aligned to one incident window?
Comparing unrelated time ranges can create false correlations. If a GPU warning occurred yesterday and an application failure occurs today, placing both on the same dashboard without time context can make them look related.
The source design uses incident timelines and time-window correlation. A good RCA workflow identifies a relevant window around the incident and compares hardware events, resource metrics, network metrics, storage metrics, container state, application state and changes within it. The precise window depends on the incident; the source does not prescribe one universal duration.
What do historical incidents contribute?
Historical incidents provide evidence about recurring patterns. The source knowledge model stores work orders, alarm events, postmortems and operational experience, which lets the platform ask whether this object has failed this way before, whether this error code has appeared before, whether the same sequence led to hardware replacement, and which remediation worked.
The source root-cause example uses recent similar hardware cases as one evidence item. History can raise or lower confidence in the current hypothesis, but it should not replace current evidence.
Why can historical similarity be misleading?
Similar symptoms can come from different causes. An application timeout can be caused by network, storage, database, compute or service overload, and a previous incident may look similar at the application layer while differing underneath.
The source design avoids relying on history alone by combining it with current topology and time-series evidence. Historical incidents answer what happened before, and current topology and metrics answer whether that explanation fits this incident now.
What do configuration relationships contribute?
Configuration relationships connect objects across physical and logical layers. The source CMDB model links accelerator card, server, rack, network, storage, container, application, model service, project and owner, which allows the RCA engine to understand shared dependencies.
If several failing workloads are bound to the same card, that relationship matters. So does one storage path that supports every affected task. Without accurate relationships, root-cause analysis turns into guesswork across disconnected monitoring systems.
How does configuration change history help?
Change history adds an important question: what changed before the failure?
The source data foundation keeps automatic configuration changes and point-in-time history, and the governance layer records before and after values. RCA can then compare the incident start with a recent firmware change, network configuration change, resource reassignment, driver update, model deployment, quota change or routing change.
A recent change is evidence without being proof, but it can narrow the investigation quickly. If symptoms begin immediately after one relevant configuration change, that change deserves attention.
How can topology and change history work together?
Topology identifies what the change could affect. Suppose a network configuration changes on Switch A at 12:01. At 12:03, three services become unhealthy, and topology shows all three depend on Switch A. That combination is much stronger than either signal alone, while a change on an unrelated switch at the same time would be weaker evidence. This is the value of combining relationship and temporal information.
What are same-source alarms?
The source Q&A lists same-source alarms as one input to root-cause confidence. Repeated hardware errors from one accelerator card can strengthen the hypothesis that the card is the problem.
The exact implementation of same-source scoring is not specified. The principle is that repeated related evidence from one object should be weighed alongside other domains, and a burst of same-source alarms means more when topology also shows that the affected workloads depend on that object.
What do rule matches contribute?
Rules represent known operational knowledge. A rule can say that an event combination is strongly associated with a known hardware condition, that a certain downstream alarm commonly follows one upstream event, or that one critical event should never be suppressed.
The source root-cause model includes rule matches as one confidence input. Rules are valuable for well-understood infrastructure failure patterns, and they should stay visible as part of the evidence.
How should confidence be presented?
Present confidence with the evidence behind it; the source explicitly rejects an unexplained score. A useful RCA result can say:
Likely root cause: GPU card 3.
Confidence: high.
Evidence: ECC count increased before container restart, all affected containers were bound to card 3, peer cards remained normal, and recent similar incidents ended in hardware repair.
An operator can act on that, which is more than a score alone allows.
How can RCA handle several plausible causes?
Keep more than one hypothesis when the evidence is inconclusive. The source does not specify a formal multi-hypothesis engine, but its requirement for confidence and itemized evidence supports this behavior.
For example, a storage bottleneck may have high evidence, network loss medium evidence, and a GPU fault low evidence. The operator can then run the verification step that best distinguishes the candidates. Do not force one high-confidence answer when the data is ambiguous.
How does business impact connect to root cause?
After the likely cause is identified, traverse relationships upward to determine the affected service. The source model connects infrastructure to tasks, model services, applications, projects and owners, which lets RCA move from a technical condition to a statement about affected training jobs, inference services, applications or business systems.
For the impact-analysis layer, how network topology and application topology help identify the business impact of infrastructure failures explains how the same graph is used upward.
How can an RCA result drive remediation?
The source incident workflow matches the root-cause result to a runbook and response recommendation. The action then follows the approved remediation level, which is automatic, semi automatic or manual, and the work order records execution and recovery.
If the remediation works, the incident outcome becomes historical evidence. If it fails, the platform learns that the hypothesis or runbook needs review.
For the automation boundary, what is the difference between automatic, semi automatic, and manual remediation in IT operations explains how risk should determine execution mode.
How should postmortems improve RCA?
Postmortems turn investigated incidents into structured historical evidence. The source automatically places reviewed postmortems into the knowledge base, so future RCA can use the accepted root cause, contributing factors, timeline, remediation and improvement actions. That is stronger than relying only on raw alarm similarity.
For that process, how incident postmortems can be generated automatically from alarms, timelines, work orders, and remediation actions explains how the evidence becomes reusable knowledge.
What data-quality problems can break RCA?
Poor identity and stale relationships can break the analysis. Examples include a container mapped to the wrong node, stale network topology, an outdated server location, a missing business owner, an unrecorded configuration change, or a historical incident tied to the wrong device.
The source repeatedly treats CMDB accuracy as a prerequisite for AIOps, so configuration discovery and change tracking count as diagnostic inputs in their own right instead of separate administrative functions.
What should a root-cause interface show?
A practical interface can show the likely root cause, confidence, evidence list, incident timeline, topology path, relevant metrics, recent changes, similar historical incidents, affected services and recommended response, and it should let the operator drill into every evidence item.
A platform example that combines topology, time-series data, historical operational knowledge, and configuration relationships is Sensaka.
If I were evaluating RCA, I wouldn't stop at whether the product can produce a root-cause label. I would ask whether it can show why that component is the leading explanation, which dependencies connect it to the symptoms, what happened first, what changed recently, and whether similar reviewed incidents support the conclusion.
Frequently Asked Questions
What evidence does the source use for root-cause confidence?
The source says confidence can combine time order, topology relationships, same-source alarms, historical cases, and rule matches, and each supporting item stays visible to the operator.
Why are configuration relationships important to root cause analysis?
They connect devices, components, workloads, applications, owners, and recent configuration state, so the platform can see shared dependencies and tell whether a recent change may be relevant.
Should root-cause analysis return only one confidence score?
No. The source explicitly says the platform should display the evidence behind the conclusion instead of presenting an unexplained score.