OpenTelemetry

OpenTelemetry (OTel) is an open-source, vendor-neutral standard for generating, collecting, and exporting telemetry — traces, metrics, and logs (with profiles emerging). It's a CNCF project formed from the merger of OpenTracing and OpenCensus, and it's supported by every major observability vendor, including Datadog, Grafana, New Relic, Honeycomb, Elastic, and the big cloud providers.

The key idea: instrument once, send anywhere. Your code and libraries emit telemetry through the OpenTelemetry API; the SDK and the Collector decide where it goes. Switching or combining backends becomes configuration, not a re-instrumentation project.

TL;DR

Quick Example

Instrument a Node.js service: auto-instrumentation plus a custom span with attributes.

A Collector that receives OTLP, batches and redacts, tail-samples, and fans out to two backends:

Core Concepts

Signals

API, SDK, and Instrumentation

The API is what libraries and applications call; it's a safe no-op if no SDK is configured. The SDK implements sampling, processing, and export. Instrumentation libraries (and zero-code agents for Java, .NET, Python, Node.js) hook into frameworks, HTTP clients, and database drivers automatically.

Traces, Spans, and Context Propagation

A trace is a tree of spans, each with a name, timing, attributes, events, and status. Context propagation passes the trace ID and parent span ID between services — in HTTP via the W3C traceparent header, in messaging via message headers — so work across services joins one trace. Baggage carries additional key-value context. See Distributed Tracing.

Resources and Semantic Conventions

A resource describes what's producing telemetry (service.name, service.version, deployment.environment.name, k8s.pod.name). Semantic conventions define standard attribute names (http.request.method, http.response.status_code, db.system), making dashboards and queries portable across services and vendors.

OTLP and the Collector

OTLP is OpenTelemetry's protocol over gRPC or HTTP. The Collector is a standalone service built from receivers (OTLP, Prometheus scrape, Kafka, host metrics, filelog), processors (batch, memory limiter, attributes, filter, tail sampling, k8s attributes), and exporters (OTLP, Prometheus, vendor-specific). Deploy it as an agent (per node or sidecar) and/or a gateway (central pool).

Sampling

Best Practices

Start With Auto-Instrumentation

It gives immediate value for HTTP, database, and messaging calls. Add manual spans and attributes for business operations afterward.

Always Set service.name and Environment

Unnamed services show up as unknown_service and break service maps and dashboards.

Route Through a Collector

Applications export to a local Collector; the Collector handles retries, batching, redaction, and backend changes without redeploying apps.

Follow Semantic Conventions

Use standard attribute names and consistent units so cross-service queries work.

Redact Sensitive Data

Strip PII and secrets from attributes and logs in the Collector before data leaves your network.

Correlate Logs With Traces

Use OTel logging bridges or inject trace_id and span_id into logs so you can pivot between signals.

Common Mistakes

Broken Context Propagation

Missing propagators, custom HTTP clients, or message queues without header propagation split one request into many disconnected traces.

High-Cardinality Metric Attributes

User IDs or full URLs on metrics explode storage costs. Keep those on spans, not metrics.

Forgetting to End Spans

Spans that never end are never exported and can leak memory. Use finally blocks or helper wrappers.

Exporting Directly to a Vendor From Every Service

It works, but changing vendors or adding redaction later means touching every service.

100% Trace Sampling at Scale

Keeping every trace in high-traffic systems is expensive and rarely useful. Use tail sampling to keep what matters.

Comparison

FAQ

What is OpenTelemetry?

An open standard and set of tools for instrumenting software to produce traces, metrics, and logs, and for collecting and exporting that telemetry to any compatible backend.

Is OpenTelemetry a monitoring tool?

No. It produces and transports telemetry. You still need backends to store and visualize it, such as Grafana Tempo and Prometheus, Jaeger, or a commercial platform.

Do I need the OpenTelemetry Collector?

Not strictly — SDKs can export directly. But the Collector is recommended for production because it centralizes batching, retries, sampling, redaction, and routing.

What's the difference between head and tail sampling?

Head sampling decides whether to keep a trace when it starts. Tail sampling decides after it completes, allowing policies like "keep all errors and slow traces."

Does OpenTelemetry replace Prometheus?

Not necessarily. OTel can produce metrics and export them to Prometheus-compatible storage; the Collector can also scrape Prometheus endpoints. They work together.

Related Topics

References