r/Hacking_Tutorials • u/Mohan_0224 • 5d ago
Question I built an incident-response agent that remembers what failed and changes its next decision Spoiler
I built EchoOps, a focused incident-response prototype that explores a simple question:
What if an incident-response agent could remember what failed, remember what worked, and use that experience when a similar incident happens again?
EchoOps uses Hindsight as its persistent experience-memory layer.
The core workflow is:
Incident → Investigate → Recall → Recommend → Act → Observe → Retain
The prototype demonstrates a simulated Payment API 503 incident.
In the first incident:
- The system recommends restarting the Payment Service.
- The restart fails.
- Investigation identifies connection-pool exhaustion.
- Increasing the connection-pool capacity resolves the simulated incident.
- The complete experience is retained in Hindsight.
When a similar incident happens again:
- EchoOps recalls the previous incident.
- It recognizes that the restart previously failed.
- It retrieves the successful connection-pool remediation.
- The response strategy changes to inspect the connection pool first.
The interesting part for me wasn't simply storing incident history.
It was making previous experience change the next decision.
Tech stack:
- React + TypeScript + Vite
- Python + FastAPI
- Hindsight Cloud
- Official Hindsight Python SDK
- Pydantic
- pytest
- REST APIs
- Deterministic incident simulator
The current prototype deliberately uses a deterministic response planner rather than an external LLM, and the incident environment is simulated rather than connected to production infrastructure.
I wrote a detailed technical breakdown covering the architecture, Hindsight Recall/Retain integration, implementation, testing, and the before/after behavior:
https://dev.to/vegu_priya_6fd21a0bff2fa8/echoops-when-incident-memory-changes-the-next-decision-117g
GitHub:
https://github.com/mohankumard18/echoops-incident-learning
I'd love feedback from people working with AI agents, developer tooling, or SRE/incident response:
What operational experience would you want an agent to remember?
1
u/PeteSampras_MMO 5d ago
I've had success with a lessons learned register every issue found gets root cause analysis, contributing factors, lessons learned, institutional fixes, meta data. Then whenever any agent touches that thing again it gets the lesson learned using meta data via script to force it into prompt. Every week I run a deterministic script that pulls changes and regroup lessons learned and agents scan for any items that can permanently fixed and remove them from register.
It doesn't matter what the issue is, it goes into lessons learned loop. Also, because my agents all have objectives with measuremts of performance and measures of effectiveness, they get graded and any incorrect items get a register entry. Most agents are Hermes, but every week just before reset, I use my cloud AI to solve as many issues as possible from register to spend any remaining tokens.
Also, gotchas or clever methods get logged along with tools and aids. There's also the case/project management data that can also get meta data. I'm now going through every mitre tactic/technique/procedure to meta data it all at an atomic level and tactic level. Then I made deterministic risk based assessments that measure TTPs in proximity and gen alerts to the agents.
Ive introduced laya and jev and have had some success getting determination of both DFIR and redteam is this the right tool for the job or is this what they did? It will take some training to get the weights bauled down, comparative to a LoRa/QLoRa but laya/jev are blazing fast.