OpenTelemetry Collector

The OpenTelemetry Collector is a standalone service that receives, processes, and exports telemetry (traces, metrics, and logs). Instead of every application sending data directly to a vendor, apps send OpenTelemetry data (usually over OTLP) to a Collector, which batches it, enriches it with metadata, filters and samples it, and forwards it to one or many backends: Prometheus, Jaeger or Tempo, Loki, Elasticsearch, Datadog, Honeycomb, or cloud observability services.

Putting a Collector in the middle decouples instrumentation from backends. Switching vendors, adding a second destination, dropping noisy data, redacting PII, or changing sampling becomes a configuration change in one place instead of a code change in every service.

TL;DR

Quick Example

A gateway Collector receiving OTLP, enriching, sampling, and exporting to multiple backends:

Core Concepts

Components

A component only runs if a pipeline references it. Processor order matters: memory_limiter goes first, and batch usually last.

Pipelines

Each pipeline handles one signal type (traces, metrics, or logs) and lists receivers, processors, and exporters. You can define several pipelines per signal (traces/internal, traces/customer) and fan out to multiple exporters. Connectors let one pipeline's output feed another, for example generating span metrics from traces for Prometheus.

Deployment Patterns

Distributions

OTTL and Transformation

The OpenTelemetry Transformation Language (used by the transform and filter processors) edits telemetry with statements like set(attributes["tenant"], resource.attributes["k8s.namespace.name"]) or drops health-check spans with IsMatch(attributes["url.path"], "/healthz"). It's how you normalize, enrich, and reduce data without touching application code.

Best Practices

Always Configure memory_limiter and batch

memory_limiter protects the Collector from OOM crashes by refusing data under pressure (clients retry); batch greatly improves export efficiency. They're the two processors every pipeline should have.

Monitor the Collector Itself

Scrape the Collector's internal metrics: accepted, refused, and dropped spans and points, exporter queue size, send failures, and memory. Alert on refused and dropped data and on growing queues, because a silently failing Collector means silently missing telemetry.

Enable Persistent Queues for Critical Paths

Exporters' sending_queue can use the file_storage extension, so buffered data survives restarts and backend outages, trading disk for durability.

Filter and Redact Early

Drop noisy telemetry (health checks, debug logs, unused metrics) and remove sensitive attributes (tokens, emails, card numbers) at the agent or gateway. That reduces cost and compliance risk. See PII handling.

Common Mistakes

Tail Sampling Without Trace-Aware Load Balancing

Spreading one trace's spans across several gateway replicas means each replica sees partial traces and makes inconsistent sampling decisions. Route by trace ID with the loadbalancing exporter.

Using the debug Exporter in Production

The debug exporter (formerly logging) at detailed verbosity prints every span to stdout, which floods logs and burns CPU. Use it only while troubleshooting.

One Giant Pipeline for Everything

Mixing signals from different tenants or sensitivity levels in one pipeline makes redaction, routing, and quotas hard. Split pipelines by purpose, or use routing connectors.

FAQ

Do I need a Collector, or can apps export directly?

Apps can export OTLP directly to many backends, which is fine for small setups. A Collector adds batching, retries, enrichment, sampling, redaction, multi-backend export, and vendor independence without app changes. Most production deployments use one.

What's the difference between agent and gateway mode?

The Collector binary is the same; the deployment differs. Agents run close to workloads (per node or pod) to collect local data and add host context. Gateways run as a central, scalable tier for cross-cutting processing like tail sampling and routing to vendors. Many deployments use both.

Can the Collector replace Prometheus scraping?

The prometheus receiver can scrape Prometheus endpoints using standard scrape configs, and export via remote write or OTLP. It's a good fit when you want one agent for all signals. Prometheus itself still offers the richer local query and alerting engine.

How do I run the Collector on Kubernetes?

Use the OpenTelemetry Operator, which manages OpenTelemetryCollector resources in deployment, DaemonSet, sidecar, or StatefulSet modes, plus Instrumentation resources for automatic SDK injection. The official Helm charts are an alternative.

Related Topics

References