Problem
Routine production failures (stuck workers, memory pressure, a bad rollout) consume on-call time, and they arrive as a storm of isolated alerts rather than a single cause.
Case study 02 · Reliability engineering
A constrained multi-agent operator that detects production anomalies, reasons about root cause, and applies reversible fixes.

Routine production failures (stuck workers, memory pressure, a bad rollout) consume on-call time, and they arrive as a storm of isolated alerts rather than a single cause.
Logs, metrics and traces are watched continuously and anomalies are grouped into incidents. A diagnosing agent correlates recent deploys, resource pressure and error signatures, then proposes the smallest change likely to fix the incident.
Walk through the pipeline: select a stage, or use the arrow keys.
Prometheus metrics, pod events and logs are collected continuously as the system's view of production.
Telemetry
Prometheus metrics, pod events and logs are collected continuously as the system's view of production.
The agent can only act through an allow-list of reversible operations (scaling, restarts, rollbacks) with blast-radius limits and a dry run first. There is no arbitrary shell access.
Every action is followed by a health check against the signals that triggered it. If the system does not recover, the change is rolled back automatically and the incident is escalated to a person.
Routine failures resolve themselves and write their own incident record (what was seen, what was tried, what changed), so people can focus on the genuinely novel problems.

Next step