Skip to content
Srikanta Sahu

Projects06 Platform engineering

Observability-as-Code Platform

Application observability defined as configuration, then provisioned through Terraform, Helm, and GitOps.

Problem

Each application was onboarded to metrics, logs, traces, dashboards, and SLOs by hand, so the standard drifted and the next service started from scratch.

Solution

Defined application observability as configuration and automatically provisioned the collectors, dashboards, alerts, and SLO resources from that definition.

Architecture

  1. App definition
  2. CI validation
  3. Terraform / Helm
  4. Telemetry
  5. Dashboards and SLOs
  6. Deploy

Technology

  • Terraform
  • Kubernetes
  • Helm
  • Prometheus
  • Grafana
  • Loki
  • OpenTelemetry
  • GitHub Actions

Engineering decisions

  • Reusable modules that still fit services with different signals.
  • Detecting drift when someone changes a dashboard outside Git.
  • Policy checks so an application cannot skip required alerts.
  • A default onboarding path that is the same for every team.
  • A GitOps workflow where Git remains the source of the live config.

Automation

A merged definition creates or updates collectors, dashboards, and alerts. Teams do not click those resources together in the UI.

Reliability

CI rejects a definition that is missing required fields. Apply is incremental, so a bad module does not rebuild the whole platform.

Security

Modules create scoped service accounts and keep datasource credentials in the secret store, not in the application repo.