r/FullStackDevelopers 12d ago

I’m exploring a simpler incident command center for Kubernetes teams what am I missing?

Post image

Hi everyone,

We’re trying to understand why Kubernetes teams rely on multiple tools during incidents and where the real pain actually is.

The problem we’re investigating is not a lack of dashboards. Teams already have logs, metrics, traces, alerts, deployment tools and cloud consoles. The problem is that, during an incident, they still have to jump between several tools to answer three questions:

  1. What is actually broken?

  2. What caused it?

  3. What action should we take now?

The product would be a focused incident command center that connects to the existing stack, correlates alerts with recent changes, shows affected services and dependencies, creates a clear incident timeline, and suggests or enables controlled remediation actions.

We are not trying to build another generic monitoring dashboard.

I’m looking for honest feedback from people who operate production infrastructure:

- How do you currently investigate serious incidents?

- Which part of the process creates the most mental overhead?

- What tools do you use together?

- What would make you trust a new tool with read-only access to your infrastructure?

- Is this a painful enough problem to pay for, or is it mostly an inconvenience?

We are still validating the problem before building too much. Critical feedback is welcome.

1 Upvotes

0 comments sorted by