Skip to content
Srikanta Sahu

Projects03 AI / SRE

AI SRE / Autonomous Incident Investigation

An investigation assistant that queries metrics, logs, cluster state, and runbooks, then drafts an explainable RCA for a human to approve.

Problem

When an incident fires, an engineer still has to query metrics, logs, cluster state, and runbooks by hand, then rebuild a timeline before they can recommend a fix. That is slow and easy to get wrong.

Solution

Built an investigation assistant that queries observability and infrastructure systems, correlates evidence, consults runbooks, constructs a timeline, and produces an explainable root-cause write-up with a remediation recommendation. A human still approves any change. The assistant does not remediate on its own.

Architecture

  1. Incident
  2. Investigation
  3. Prometheus
  4. Loki
  5. Kubernetes
  6. Runbook RAG
  7. Timeline
  8. Root cause
  9. Human approval

Technology

  • Prometheus
  • Loki
  • OpenTelemetry
  • Kubernetes
  • Python
  • LLM
  • RAG
  • MCP
  • n8n

Engineering decisions

  • Tool access with narrow permissions so the assistant can query without changing the system.
  • Evidence-based RCA: every claim has to point at a metric, log, event, or runbook passage.
  • Correlating metrics, logs, and cluster events onto one timeline.
  • Keeping a human in the loop so a recommendation never becomes an unreviewed change.

Automation

The first investigation pass — gathering signals, assembling a timeline, and drafting the RCA — no longer starts as a blank page.

Reliability

If a tool call fails, the write-up says which evidence is missing instead of guessing. The path stops at a recommendation.

Security

Credentials stay on the server. The assistant can read the signals it is allowed to see and cannot apply a change without approval.

AI layer

RAG retrieves runbooks and postmortems. MCP exposes metrics, logs, and cluster state as tools. The model proposes; a person decides.