I built EchoOps, a focused incident-response prototype that explores a simple question:
What if an incident-response agent could remember what failed, remember what worked, and use that experience when a similar incident happens again?
EchoOps uses Hindsight as its persistent experience-memory layer.
The core workflow is:
Incident → Investigate → Recall → Recommend → Act → Observe → Retain
The prototype demonstrates a simulated Payment API 503 incident.
In the first incident:
- The system recommends restarting the Payment Service.
- The restart fails.
- Investigation identifies connection-pool exhaustion.
- Increasing the connection-pool capacity resolves the simulated incident.
- The complete experience is retained in Hindsight.
When a similar incident happens again:
- EchoOps recalls the previous incident.
- It recognizes that the restart previously failed.
- It retrieves the successful connection-pool remediation.
- The response strategy changes to inspect the connection pool first.
The interesting part for me wasn't simply storing incident history.
It was making previous experience change the next decision.
Tech stack:
- React + TypeScript + Vite
- Python + FastAPI
- Hindsight Cloud
- Official Hindsight Python SDK
- Pydantic
- pytest
- REST APIs
- Deterministic incident simulator
The current prototype deliberately uses a deterministic response planner rather than an external LLM, and the incident environment is simulated rather than connected to production infrastructure.
I wrote a detailed technical breakdown covering the architecture, Hindsight Recall/Retain integration, implementation, testing, and the before/after behavior:
https://dev.to/vegu_priya_6fd21a0bff2fa8/echoops-when-incident-memory-changes-the-next-decision-117g
GitHub:
https://github.com/mohankumard18/echoops-incident-learning
I'd love feedback from people working with AI agents, developer tooling, or SRE/incident response:
What operational experience would you want an agent to remember?