Trace Sampling in OpenTelemetry
Recording every span of every request is rarely affordable. A service handling 5,000 requests per second, each producing 20 spans, generates 100,000 spans per second, and storage and ingestion bills scale linearly. Sampling keeps a representative or targeted subset of traces while discarding the rest, which cuts cost and overhead dramatically.
The challenge is keeping the interesting traces: errors, slow requests, rare paths. OpenTelemetry supports head sampling (decide at the start of a trace, in the SDK) and tail sampling (decide after the trace completes, in the Collector). Most production setups combine them, and compute metrics before sampling so dashboards stay accurate.
TL;DR
- Head sampling decides when the root span starts. It's cheap and simple, but blind to outcome (errors and latency aren't known yet).
ParentBased(TraceIdRatioBased(p))is the standard SDK sampler: roots sample at probability p, and children follow the parent's decision.- Tail sampling (Collector
tail_samplingprocessor) decides after seeing the whole trace, so it can keep all errors and slow traces, plus a baseline percentage. - Tail sampling needs all spans of a trace in one Collector instance: route by trace ID with the
loadbalancingexporter. - Derive RED metrics from unsampled data (span metrics connector or app metrics), so sampling doesn't distort rates.
- Record sampling probability so backends can extrapolate counts correctly.
Quick Example
Head sampling in the SDK via environment variables:
Tail sampling in a gateway Collector:
(The forward connector links the two trace pipelines.) Every error and every trace over 800 ms is kept, plus 2% of the rest, while request-rate and latency metrics reflect 100% of traffic.
Core Concepts
Head Sampling
The sampler runs when a span starts, deciding RECORD_AND_SAMPLE, RECORD_ONLY, or DROP. Built-in samplers:
Because TraceIdRatioBased uses the trace ID, every service with the same ratio makes the same decision for a trace, and ParentBased makes downstream services honor the upstream decision carried in traceparent (see context propagation). The result is complete traces or none.
- Pros: negligible overhead, since unsampled spans are never exported; simple; no infrastructure.
- Cons: you can't prefer errors or slow requests, because the decision happens before anything goes wrong.
Tail Sampling
The Collector buffers all spans of a trace for decision_wait, then evaluates policies:
status_code(errors),latency(slow traces),string_attribute/numeric_attribute(specific routes, tenants, versions),probabilistic(baseline),rate_limiting(cap spans per second),span_count,ottl_condition, andand/compositecombinations.
A trace is kept if any policy says so (composite policies allow budgets per policy).
- Pros: keeps what matters, with the full picture of errors and outliers.
- Cons: every span must be sent to the Collector (network and CPU cost upstream), memory to buffer traces, a fixed decision delay, and it requires trace-ID-aware routing when Collectors scale horizontally.
Routing for Tail Sampling
With several gateway replicas, spans of one trace may hit different replicas. Put a load-balancing tier in front, where agent Collectors use the loadbalancing exporter with routing_key: traceID, so each trace's spans reach the same sampler instance.
Keeping Metrics Accurate
Sampled traces give skewed counts: 2% of normal traffic plus 100% of errors makes the error rate look huge. Solutions:
- Compute RED metrics before sampling: the Collector's
spanmetricsconnector, application metrics via OpenTelemetry instrumentation, or Prometheus client libraries. - Use sampling-aware backends that read the sampling probability (the
ththreshold intracestatefor consistent probability sampling) and weight counts accordingly.
Consistent Probability Sampling
The OpenTelemetry specification defines consistent probability sampling, where the sampling threshold travels in tracestate (ot=th:…), so downstream services and backends know each span's effective probability. It enables accurate extrapolation and mixing different sampling rates across services.
Choosing a Strategy
Best Practices
Sample Traces, Not Metrics
Keep metrics unsampled so SLOs, dashboards, and alerts reflect reality. Use traces for investigation, where a well-chosen subset suffices.
Always Keep Errors
Whatever the strategy, make sure failed requests are retained. Tail sampling does this directly; with head-only sampling, rely on logs and error tracking (Sentry) for full error coverage.
Size Tail Sampling Resources Deliberately
Memory is roughly spans per second × decision_wait × average span size. Monitor dropped traces (otelcol_processor_tail_sampling_* metrics) and sampling decision latency, and scale horizontally with trace-ID load balancing.
Communicate the Sampling Rate
Engineers looking at traces should know whether a missing trace means "didn't happen" or "wasn't sampled". Document rates and record them as resource or span attributes.
Common Mistakes
Non-Parent-Based Samplers Downstream
If service B uses TraceIdRatioBased(0.1) without ParentBased, it may drop spans from traces service A decided to keep, and vice versa. You end up with fragments. Use parent-based samplers everywhere except the true entry points.
Computing Error Rates From Tail-Sampled Traces
A tail sampler that keeps all errors but 2% of successes inflates error ratios about 50×. Compute rates from unsampled metrics.
Tail Sampling Without Enough decision_wait
Long-running traces (async processing, slow downstream calls) may still be receiving spans when the decision is made, so late spans are dropped or become orphaned. Tune decision_wait to your trace durations, or accept partial traces for long workflows.
FAQ
Head or tail sampling?
Head sampling is simple and cheap and works well when you mostly need representative traces. Tail sampling is better when you must keep every error or slow request, at the cost of Collector infrastructure and memory. Many teams start with head sampling and add tail sampling as volume and needs grow.
What sampling rate should I use?
Enough to see typical requests for every important endpoint regularly: often 1–10% for high-traffic services, more for low-traffic ones, plus 100% of errors and slow traces with tail sampling. Let budget and query needs guide it, and remember that metrics shouldn't depend on it.
Does sampling reduce instrumentation overhead?
Head sampling does: unsampled spans are cheap no-ops and never exported. Tail sampling doesn't reduce application overhead, because all spans are exported to the Collector. It reduces storage and backend costs.
Can I force a specific request to be traced?
Yes, through several mechanisms: send a request with a sampled traceparent (parent-based samplers honor it), use a tail-sampling attribute policy (such as debug=true baggage copied to an attribute), or use vendor-specific force-sample headers. Restrict this at the edge so external clients can't force-sample everything.
Related Topics
- OpenTelemetry — The project overview
- OpenTelemetry Collector — Where tail sampling runs
- OpenTelemetry Context Propagation — How sampling decisions travel
- Distributed Tracing — Tracing fundamentals
- Cloud Costs — Observability as a cost center
- SLOs — Why metrics should stay unsampled