Trace Sampling in OpenTelemetry

Recording every span of every request is rarely affordable. A service handling 5,000 requests per second, each producing 20 spans, generates 100,000 spans per second, and storage and ingestion bills scale linearly. Sampling keeps a representative or targeted subset of traces while discarding the rest, which cuts cost and overhead dramatically.

The challenge is keeping the interesting traces: errors, slow requests, rare paths. OpenTelemetry supports head sampling (decide at the start of a trace, in the SDK) and tail sampling (decide after the trace completes, in the Collector). Most production setups combine them, and compute metrics before sampling so dashboards stay accurate.

TL;DR

Quick Example

Head sampling in the SDK via environment variables:

Tail sampling in a gateway Collector:

(The forward connector links the two trace pipelines.) Every error and every trace over 800 ms is kept, plus 2% of the rest, while request-rate and latency metrics reflect 100% of traffic.

Core Concepts

Head Sampling

The sampler runs when a span starts, deciding RECORD_AND_SAMPLE, RECORD_ONLY, or DROP. Built-in samplers:

Because TraceIdRatioBased uses the trace ID, every service with the same ratio makes the same decision for a trace, and ParentBased makes downstream services honor the upstream decision carried in traceparent (see context propagation). The result is complete traces or none.

Tail Sampling

The Collector buffers all spans of a trace for decision_wait, then evaluates policies:

A trace is kept if any policy says so (composite policies allow budgets per policy).

Routing for Tail Sampling

With several gateway replicas, spans of one trace may hit different replicas. Put a load-balancing tier in front, where agent Collectors use the loadbalancing exporter with routing_key: traceID, so each trace's spans reach the same sampler instance.

Keeping Metrics Accurate

Sampled traces give skewed counts: 2% of normal traffic plus 100% of errors makes the error rate look huge. Solutions:

Consistent Probability Sampling

The OpenTelemetry specification defines consistent probability sampling, where the sampling threshold travels in tracestate (ot=th:…), so downstream services and backends know each span's effective probability. It enables accurate extrapolation and mixing different sampling rates across services.

Choosing a Strategy

Best Practices

Sample Traces, Not Metrics

Keep metrics unsampled so SLOs, dashboards, and alerts reflect reality. Use traces for investigation, where a well-chosen subset suffices.

Always Keep Errors

Whatever the strategy, make sure failed requests are retained. Tail sampling does this directly; with head-only sampling, rely on logs and error tracking (Sentry) for full error coverage.

Size Tail Sampling Resources Deliberately

Memory is roughly spans per second × decision_wait × average span size. Monitor dropped traces (otelcol_processor_tail_sampling_* metrics) and sampling decision latency, and scale horizontally with trace-ID load balancing.

Communicate the Sampling Rate

Engineers looking at traces should know whether a missing trace means "didn't happen" or "wasn't sampled". Document rates and record them as resource or span attributes.

Common Mistakes

Non-Parent-Based Samplers Downstream

If service B uses TraceIdRatioBased(0.1) without ParentBased, it may drop spans from traces service A decided to keep, and vice versa. You end up with fragments. Use parent-based samplers everywhere except the true entry points.

Computing Error Rates From Tail-Sampled Traces

A tail sampler that keeps all errors but 2% of successes inflates error ratios about 50×. Compute rates from unsampled metrics.

Tail Sampling Without Enough decision_wait

Long-running traces (async processing, slow downstream calls) may still be receiving spans when the decision is made, so late spans are dropped or become orphaned. Tune decision_wait to your trace durations, or accept partial traces for long workflows.

FAQ

Head or tail sampling?

Head sampling is simple and cheap and works well when you mostly need representative traces. Tail sampling is better when you must keep every error or slow request, at the cost of Collector infrastructure and memory. Many teams start with head sampling and add tail sampling as volume and needs grow.

What sampling rate should I use?

Enough to see typical requests for every important endpoint regularly: often 1–10% for high-traffic services, more for low-traffic ones, plus 100% of errors and slow traces with tail sampling. Let budget and query needs guide it, and remember that metrics shouldn't depend on it.

Does sampling reduce instrumentation overhead?

Head sampling does: unsampled spans are cheap no-ops and never exported. Tail sampling doesn't reduce application overhead, because all spans are exported to the Collector. It reduces storage and backend costs.

Can I force a specific request to be traced?

Yes, through several mechanisms: send a request with a sampled traceparent (parent-based samplers honor it), use a tail-sampling attribute policy (such as debug=true baggage copied to an attribute), or use vendor-specific force-sample headers. Restrict this at the edge so external clients can't force-sample everything.

Related Topics

References