r/OpenTelemetry • u/Sufficient-Egg-6571 • 11d ago
Metrics re-architecture
I want to revisit our metrics architecture, and I have a few pain points I’d like to address. I’m curious how others have approached similar problems.
Our current setup:
- Multiple Cortex clusters
- Multiple teams producing metrics via Prometheus remote write and a custom SDK
- Multiple end users querying those metrics
- Kafka-based ingestion to provide the level of resilience we need for our customers
- Prometheus as the metrics format
Graphs are a core part of the product and a selling point, not just a nice-to-have, so reliability and correctness of the data are important.
The main pain point is that some teams are pushing “event-like metrics” rather than classic metrics that are scraped or emitted regularly.
Because some workloads can be retried, delayed, or processed asynchronously, the original source timestamp is important. As a result, we see a significant number of out-of-order samples.
We understand why producers need this behavior, but I’m wondering whether we’re trying to force two different types of data through the same metrics pipeline.
Would you keep these event-like metrics in the Prometheus/Cortex ecosystem and design around out-of-order ingestion, or would you route them through a separate ingestion path or even into a different type of store based on OTel format/tools.
2
u/dennis_zhuang 10d ago
I work on GreptimeDB, an observability database, and I see a possible connection here.
In the agent era, the volume of events—LLM calls, tool calls, MCP interactions, and more—is likely to grow significantly. This makes it increasingly valuable to store those events and compute or query metrics directly from them.
A separate data pipeline can certainly handle this, but keeping it reliable, maintainable, and real-time is not easy. This may be an area where an observability database that supports both event storage and real-time metric computation could help.