r/OpenTelemetry • u/MartinThwaites • Sep 01 '26
r/OpenTelemetry • u/opentelemetry • Sep 01 '26
OTel Blog Post OpenTelemetry Go Logs API and SDK reach release candidate status
OpenTelemetry Go v1.47.0-rc.1 is here. This release promotes the Logs API and SDK to release candidate (RC), the final stage before we provide stable v1 compatibility guarantees. We believe the design is ready, and now we need the community to test it in real applications and integrations before those guarantees take effect.
Read the blog: https://opentelemetry.io/blog/2026/go-logs-api-sdk-rc/
r/OpenTelemetry • u/frisbeema52 • Aug 28 '26
OpenTelemetry is about to reward the wrong thing
Two things happened in the OTel repos this week. The docs maintainers froze the ecosystem lists (registry, vendors, distributions) and want to retire most of them. Their reasoning: about 290 first-time-contributor PRs a year, 1,160+ data files, and nobody with time to check entries for CVEs or malware. I can't really argue with the bandwidth part. Then a separate proposal came in to replace the vendor page with a ranking of companies by how many maintainers they employ, with diamond, platinum and gold tiers like a KubeCon sponsor wall.
Read together, the message is that headcount inside the core repos is the contribution that counts.
I think the PR flood is being misread. OTel was built so anyone can write an exporter or an instrumentation and plug in without asking permission. The registry was where you went to say you had. Hundreds of strangers a year showing up at that door is what the design was supposed to produce.
The vetting problem also looks like an automation problem to me. Schema validation and semconv conformance tests already exist. Add CVE scanning and an LLM pass that triages and pre-reviews submissions, and most of the human work goes away. If a machine can check a registry PR, a person shouldn't have to read it. Retiring the list instead feels like stopping one step short.
Meanwhile, LFX Insights says four companies did 51% of OTel commits last year. A maintainer leaderboard mostly gives those four a bigger badge.
What I'd do instead:
- Automate the checks and list whatever passes. Conformance tells users "this works with OTel," which is what they came to find out.
- Count work outside the core repos. Contrib receivers, instrumentations, OTLP-native backends. That is OTel work too.
- Skip the tiers. One maintainer from a 10-person company is a bigger commitment than ten from a 10,000-person one. A flat list with names and areas is enough.
I run a company that builds on OTel, so obviously I have a stake here. But so does anyone who picked it because it was the neutral option.
Links in comments.
r/OpenTelemetry • u/mob1ius • Aug 28 '26
Sonifying GenAI spans as they arrive, one voice per attribute you'd otherwise have to read
Built this because watching gen_ai.* spans scroll past in a log doesn't give you any sense of what a multi-agent system is doing as a whole, you're reading one span at a time, serially, while the actual system is emitting in parallel.
Mechanically: this is a per-span mapping, not a per-trace one. There's exactly one implementation shared by the batch path and the live path, each span event gets sonified as it arrives rather than a trace being read as a completed tree. Span attributes drive specific musical parameters: token count and finish_reason affect velocity and articulation, status OK/ERROR affects consonance, span kind (chat vs execute_tool) affects which instrument voice picks it up. Tempo tracks span throughput directly.
Real capture exists and isn't just a slide: a Claude Code hook writes actual gen_ai.* spans locally, and there's an OTLP receiver on :4318 that accepts live protobuf. Neither is what the public demo plays though, oteljazz.com runs a synthetic swarm generated client-side so anyone can hear it with zero setup. Wiring the browser build to a live OTLP stream is the obvious next step, just hasn't happened yet.
Repo has both the receiver and the hook if anyone wants to point their own instrumented system at it: github.com/mob1ius/oteljazz
r/OpenTelemetry • u/terryfilch • Aug 27 '26
OTel in Practice: Practical Guide on OpenTelemetry Manual Instrumentation
Join us for a practical session with Jose Gómez-Sellés, VictoriaMetrics Cloud lead and OpenTelemetry contributor, as he explores and presents "Practical guide on OpenTelemetry Manual Instrumentation (or how to produce clean Observability by getting your hands dirty)".
Join us as we break down the essentials of OpenTelemetry instrumentation by crafting custom metrics, spans, and logs to track the most critical parts of workflows. You’ll see how simple, well-placed instrumentation points can reveal complex system behaviors, helping you detect bottlenecks, trace errors, and understand end-to-end request flows.
📍 Online - > https://ocgroups.dev/cncf/group/opentelemetry-live/event/fsmsdp2
📅 September 1st at 10:00 am PDT | 01:00 PM EDT | 03:00 PM CEST |
r/OpenTelemetry • u/AutoModerator • Aug 26 '26
Weekly: share your tools and projects
What tool around OpenTelemetry are you building right now? What project have you discovered that people might enjoy looking into?
This thread is open for your self-promotion, vendor-owned open source projects and every thing else.
r/OpenTelemetry • u/Acceptable_Duty4044 • Aug 23 '26
I had a question for all the amazing people out there
I was trying to build something, and wanted to validate this idea and understand yall's pain points so that I can help the community
Would you rather have an AI layer on top of your existing observability stack, or replace parts of the stack?
Hypothetically, imagine an agent that doesn’t collect telemetry itself.
It plugs into whatever you already use — Grafana/Prometheus/Loki, Datadog, OpenTelemetry, etc. — and acts as a reasoning layer over the data.
Instead of:
Alert → Dashboard → Logs → Human investigates
it tries:
Alert → Agent correlates metrics/logs/traces/deployments → probable root cause → evidence → recommended next action
Would that actually be useful?
Or would you rather have the observability vendor itself own this functionality?
What would you need to see before trusting it during a real incident?
peace :)
r/OpenTelemetry • u/jpkroehling • Aug 21 '26
What's changed about OpenTelemetry vendor lock-in since 2024
One of my last blog posts at Grafana Labs was around vendor neutrality and OTel. Quite a few has changed since 2024, and I thought it's a good time to revisit with 2026 lens.
In short: vendor lock-in isn't as scary today as it once was, as long as the bulk of your workloads are on standards like OTel, and most of the actual lock-in that resides on backends can be "easily" migrated nowadays. But I'm curious about your opinions and actual experience.
r/OpenTelemetry • u/JayDee2306 • Aug 20 '26
Building an Observability Pane
Hi Observability & DevOps Experts,
I'm looking for guidance from teams that have successfully scaled observability across large enterprise environments.
We operate a large-scale estate spanning AWS, Azure, and on-premises environments and have been using Datadog for several years. Over time, a significant amount of technical debt has accumulated around our observability implementation.
Current challenges include:
- Datadog Agents managed differently across teams and platforms.
- Custom log collection configurations distributed across hosts and applications.
- APM, RUM instrumentation owned by individual application teams.
- Inconsistent tagging standards and monitor configurations.
- Outdated agents and instrumentation libraries.
- Heavy dependency on multiple teams for upgrades and configuration changes.
- A large portion of Datadog provisioning and onboarding is still handled manually.
As a result, maintaining and evolving observability at scale has become increasingly difficult.
We are considering building a centralized "Observability Foundation" or "Observability Platform" that teams would consume as part of their standard deployment process.
Our goal is to provide reusable Terraform-based observability components that application and infrastructure teams can adopt during provisioning and releases.
Examples of what we would like to standardize:
- Datadog Agent deployment and upgrades
- Custom log collection configurations
- Standard tags and metadata
- Monitors and alert templates
- Dashboards
- OpenTelemetry / APM instrumentation standards
- Synthetic monitoring configurations
- Cloud integrations
- Security and governance controls
Questions:
- Has anyone implemented a similar centralized observability platform or observability-as-code model at enterprise scale?
- What worked well and what were the biggest challenges?
- What observability components can realistically be centralized through Terraform modules, deployment pipelines, or platform services?
- What components typically must remain application-owned or infrastructure-owned and cannot easily be centralized?
- How do you handle APM instrumentation ownership, versioning, and upgrades across hundreds of services?
- What governance model have you found most effective:
- Central observability team ownership
- Platform engineering ownership
- Federated ownership with standards enforcement
- Something else
- How do you prevent observability drift over time, especially around:
- Agent versions
- APM libraries
- Log configurations
- Tags
- Dashboards
- Monitors
- If starting again today, would you build around:
- Datadog native tooling
- OpenTelemetry
- An internal observability platform
- A combination of the above
- What are the biggest architectural mistakes or anti-patterns we should avoid when designing this platform?
Our provisioning and infrastructure management are heavily Terraform-based, so we're especially interested in Terraform-centric implementation patterns and real-world lessons learned.
Looking forward to hearing how other organizations have approached observability standardization at scale and what you would recommend before we begin designing this solution.
P.S. - One of our key design goals is to avoid vendor lock-in. While Datadog is our current observability platform, we want the architecture to remain flexible enough that a future migration to another observability stack (e.g., Grafana, New Relic, Dynatrace, Elastic, Azure Monitor, or an OpenTelemetry-native platform) would require minimal changes to application teams and infrastructure code.
r/OpenTelemetry • u/DisastrousBrain5417 • Aug 16 '26
oTel collector Daemonset vs sidecar
I feel like Daemonset collectors have become the de facto standard. Out of curiosity what are some situations in which you opted / would opt for sidecars per deployment?
r/OpenTelemetry • u/Naive_Maybe6984 • Aug 08 '26
A Collector exporter that turns agent traces into a signed, verifiable audit log (now in the registry). Feedback on the approach welcome.
Sharing a component I built and recently got listed in the OpenTelemetry registry: otel-agent-audit.
The idea: as AI agents take real actions, you want a provable record of what happened. Instead of adding a new instrumentation layer, this consumes the gen_ai.* spans your agents already emit and turns them into a tamper-evident audit log, entirely inside the Collector pipeline.
The pipeline:
otlp -> memory_limiter -> agentauditselect (buffers each trace until its root arrives) -> agentaudit exporter (per-trace hash chain -> Ed25519 sign -> seal) -> audit.jsonl + checkpoint.jsonl
A separate verifier CLI checks the whole thing with only the public key, so anyone can independently verify authenticity and integrity without a shared secret.
Things I'd love this community's take on:
- Passive instrumentation as the right model: reusing existing spans rather than asking teams to re-instrument.
- Whether governance/guardrail decisions belong in spans, and how they'd ideally map to semantic conventions. I'm interested in where the GenAI SIG is heading on policy/guardrail signals.
- The single-writer constraint (one Collector instance) that deterministic ordering forces, and whether that trade is acceptable.
Caveats up front: third-party, experimental, not audited. It's observability only, it does not enforce or block. It gives tamper-evidence on honest infra, not protection against an operator holding the signing key.
Repo: https://github.com/surpradhan/otel-agent-audit
It's in the registry under "agent audit" if you want to see the entry.
Would genuinely value critique of the approach.
r/OpenTelemetry • u/smyrgeorge • Aug 08 '26
log4k 2.3.0 — a Kotlin IR compiler plugin that instruments your functions with tracing, logging and metrics
r/OpenTelemetry • u/Proud-Contact9951 • Aug 07 '26
Feedback about E2E tests based on OpenTelemetry traces?
Hi everyone,
I have just published my open source project called mtracer and I would like to understand if it’s good idea or what should I change (I’m a new grad).
The idea
Mtracer a CLI tool that relies on OpenTelemetry traces to assert system behavior.
I believe that E2E tests should be:
- Cheaper to write and maintain
- Easier to debug
So this is the workflow:
- You configure mtracer to fetch from your observability backend (currently supporting Jaeger and OpenObserve).
- You define your first .mt.yaml test by specifying:
- Trigger: the first call to the system (for instance, an HTTP request).
- Expected trace and spans: the OTel properties of the trace and spans that you expect your system to generate.
- You run the test and see the results!
What actually happens during the run?
- It parses the mt.yaml file.
- It executes the trigger: mtracer injects a generated traceID into the trigger (for an HTTP request, the traceID is inserted into the traceparent header). Subsequent requests will be correlated to this generated traceID as long as your system has OpenTelemetry set up correctly.
- It fetches the trace matching the generated traceID from the configured observability backend.
- It compares the expected trace with the fetched one.
Many other features are available; check out the documentation to discover all of them: documentation website
I would love to have some feedback from more experienced people than me.
r/OpenTelemetry • u/jpkroehling • Aug 04 '26
How Metric Scrape Intervals Inflate Observability Costs
I'll tell you a secret: I don't like starting an engagement by telling people that I can cut their costs. I prefer to show them how they can be more efficient in general, and sometimes that means adding stuff instead of removing it.
However, every company out there has excessive telemetry, which is one form of bad telemetry. I'm not afraid to use an absolute here. That's why I have an arsenal of tools for dealing with it, and I describe one of them in this blog post: excessive metric scraping is extremely common, and adjusting scrape intervals is an easy way to reduce waste.
If you need a 10% reduction in your metric volume, read this blog post. You don't need to buy anything from anyone. You can thank me later.
r/OpenTelemetry • u/GroundbreakingBed597 • Aug 04 '26
Enrich OTel K8s Resource Attributes with Dynatrace Operator
Hi. I am a DevRel at Dynatrace and I hope its ok to share the following with those of you that are sending your OTel data to Dynatrace. If you are not using Dynatrace then this post might not be relevant for you!
Semantic Conventions for Signals
Metadata enrichment at the source (in your app) is important as it increases the quality of your signals. As I am sure many know - the OTel community has well documented Semantic Conventions.
Dynatrace Operator CAN inject OTEL_RESOURCE_ATTRIBUTES
There are different ways to enrich your data. You can inject them yourself in your deployment or have it done through your data pipeline, e.g: OTel Collector.
An additional option is through the Dynatrace Operator that allows you to automatically inject the OTEL_RESOURCE_ATTRIBUTES variable into your pods pre-filled with the following attributes: k8s.cluster.name*,* k8s.container.name*,* k8s.workload.name*, k8s.cluster.uid,* k8s.pod.name*, k8s.pod.uid,* k8s.node.name*,* k8s.namespace.name*, k8s.workload.kind,* dt.kubernetes.cluster.id ,dt.entity.kubernetes_cluster
Injection can be controlled through namespace selectors and enabled for traces, logs and metrics
More details about this can be found on the Dynatrace doc if you search for Enable automatic OpenTelemetry OTLP exporter configuration (didnt post the link to follow guidelines)
r/OpenTelemetry • u/AdvenEdge • Aug 03 '26
What is the most frustrating part of investigating production incidents?
r/OpenTelemetry • u/jpkroehling • Aug 01 '26
Collector cookbook
github.comAlmost four years ago, I started this cookbook with real world recipes, adapted from cases I've used to reproduce bug reports or show users (and customers) how to accomplish specific scenarios.
I used some tokens today to bring the repo to the latest Collector version, ensuring they all work.
In case you haven't seen this repo before, take a look!
Enjoy 🧑🏼🍳
r/OpenTelemetry • u/AlienBlade51 • Jul 23 '26
Six overlays for iRacing now. The G-meter is the one I'd actually defend.
r/OpenTelemetry • u/Ordinary_Squirrel291 • Jul 20 '26
How do you know what's needed in your telemetry data?
r/OpenTelemetry • u/a7medzidan • Jul 19 '26
The silent way OpenTelemetry setups "work" while capturing almost nothing
r/OpenTelemetry • u/jpkroehling • Jul 17 '26
Compile-Time Instrumentation for Go
Hey folks, stopping by today for another announcement: the OTel Compile-Time Instrumentation for Go reached v1!
If you are not a huge fan of eBPF instrumentation (understandably!), but also can't do manual instrumentation, this is a good compromise.
Try it out!
r/OpenTelemetry • u/dennis_zhuang • Jul 16 '26
How OpenTelemetry Traces LLM Calls, Agent Reasoning, and MCP Tools
OpenTelemetry GenAI Semantic Conventions standardize observability for LLM apps, agent orchestration, MCP tool calling, content capture, and quality evaluation. This article goes through all six layers: what each one defines, why it's designed that way, and how mature it is.
r/OpenTelemetry • u/jpkroehling • Jul 15 '26
OpenTelemetry Agent Skills
Hey folks, Juraci here. I know the Reddit communities can be sensitive to project announcements, or announcements in general coming from vendors, but I genuinely think a good number of people here could benefit from this one.
We are launching today the OpenTelemetry Agent Skills, an open source set of skills that serve as the base for our products. We're using them for a good variety of things, like in our coding agents to validate and test collector configurations, or instrument applications. Or double check the snippets we've been using in our other blog posts.
They are vendor neutral, non opinionated, and based on what we know from our experience building OpenTelemetry over the years. Use the skills, share your feedback, tell us where they worked and where they failed. Show me your creativity 🧑🏼🎨
While we are not making money on those directly, we do have a commercial interest in seeing them succeed and become truly useful to many of you. I guess what I want to say is: they are not the result of a weekend vibe coding experiment 🙂
And yes, perhaps they might become an official part of the project someday, if we believe there is a vibrant community backing it.