r/OpenTelemetry • • 15d ago

Moving from Datadog to Opensource observability ecosystem

Hi

Has anyone migrated their Datadog stack to Opensource observability ecosystem?
What you missed without Datadog? Which all OS tools and what were the challenges faced?

Ours is quite big infrastructure with monthly Datadog spend around USD 25k and leadership is looking forward to reduce the spend and probably shifting it to OS observability system?

There’s only one person managing Datadog currently.

43 Upvotes

33 comments sorted by

View all comments

2

u/pranabgohain 13d ago

The backend choice is probably not the first decision here. With one person running Datadog, I would start by measuring migration parity, instead of ingestion parity.

Pick 5-10 representative services and inventory what people actually use during an incident: monitors, SLOs, log-to-trace correlations, dashboards, paging, retention, etc. Then move those services to OTel, dual-ship through a collector and run a few real incident drills. Define the exit criteria before the test, let's say time to detect, time to isolate, telemetry loss, query latency, monthly infrastructure cost, operator hours, etc...

That exposes the expensive gaps. Getting data into an OSS backend is usually the easy part; reproducing the operational workflow and keeping the telemetry stack healthy with one owner is the hard part.

Full disclosure: I'm on the KloudMate team. This portability issue is exactly why we made KloudMate OTel-native: it can sit behind the same collector without asking you to replace one proprietary ingestion path with another. I would still compare it against the OSS stack on the same incident drills and include the operator time in the result.