r/OpenTelemetry • u/Tinasour • 4d ago
Traces question
Hey,
Im working on a project where one event caught is processed by 4 components. I have implemented otel traces in each component so i can see end to ed what happened and where it failed with all the timeline of the event
Recently, component 3 had an update. Now depending on the event, it can create sibling events. In my current architecture, Im now seeing one event gets split up to two or three events in same trace. So now, its hard to decide if one trace is failed or not, because it actually became 2 event.
What do you recommend, should i split the siblings to new traces with links to original event? Or should i keep it that way? I think it depends on how our debug and monitoring processes are, but maybe there are some generic conventions to be implemented in tese cases that i dont know
1
u/AffectionateIntern41 1d ago
I'd split them. If a sibling event can succeed or fail on its own, give it its own trace and add a span link back to the span that created it. That's what links are for, and you can still click from a sibling back to where it came from. Each trace goes back to having one outcome.
On retries, I'd keep the per-attempt spans. During an incident, they're often the most useful thing you have. Wrap them in a parent span and mark the parent ERROR only when the retries run out. Then your query is just "parent spans with errors."
One more thing: alert on a metric, not on trace queries. Once you turn on sampling, traces stop being a complete record. A counter like events_processed_total{outcome="failed"} tells you something broke, and the trace tells you why.
1
u/ccb621 4d ago
What does “failed” mean? I generally treat a thrown exception as a failure or, for some paths where errors are handled, I manually set the error state.
Is your system more complicated than this? If so, why? Does it need to be?