OpenTelemetry Collector
The OpenTelemetry Collector is a standalone service that receives, processes, and exports telemetry (traces, metrics, and logs). Instead of every application sending data directly to a vendor, apps send OpenTelemetry data (usually over OTLP) to a Collector, which batches it, enriches it with metadata, filters and samples it, and forwards it to one or many backends: Prometheus, Jaeger or Tempo, Loki, Elasticsearch, Datadog, Honeycomb, or cloud observability services.
Putting a Collector in the middle decouples instrumentation from backends. Switching vendors, adding a second destination, dropping noisy data, redacting PII, or changing sampling becomes a configuration change in one place instead of a code change in every service.
TL;DR
- The Collector is built from receivers (ingest), processors (transform), and exporters (send), wired into pipelines per signal type.
- OTLP (gRPC 4317, HTTP 4318) is the native protocol; receivers also accept Prometheus, Jaeger, Zipkin, Fluent Forward, host metrics, and many more.
- Deploy as an agent (DaemonSet or sidecar next to apps) and/or a gateway (a central, scalable tier).
- Always include
memory_limiterandbatchprocessors; addk8sattributesandresourceprocessors for context. - Tail sampling in a gateway keeps complete traces for errors and slow requests while dropping routine ones.
- Use the contrib or a custom-built distribution (OCB), and the OpenTelemetry Operator on Kubernetes.
Quick Example
A gateway Collector receiving OTLP, enriching, sampling, and exporting to multiple backends:
Core Concepts
Components
A component only runs if a pipeline references it. Processor order matters: memory_limiter goes first, and batch usually last.
Pipelines
Each pipeline handles one signal type (traces, metrics, or logs) and lists receivers, processors, and exporters. You can define several pipelines per signal (traces/internal, traces/customer) and fan out to multiple exporters. Connectors let one pipeline's output feed another, for example generating span metrics from traces for Prometheus.
Deployment Patterns
- Agent: a Collector per node (DaemonSet) or per pod (sidecar), close to applications. It gives low-latency local export, adds host and Kubernetes metadata, collects node logs and host metrics, and buffers during backend blips.
- Gateway: a horizontally scaled Collector deployment behind a load balancer. It centralizes credentials, tail sampling, routing, redaction, and vendor exporters.
- Agent → gateway is the common production shape. Tail sampling requires all spans of a trace to reach the same gateway instance, so use the
loadbalancingexporter in the agent tier to route by trace ID.
Distributions
- otelcol (core): a minimal set of components.
- otelcol-contrib: hundreds of community components, the most commonly used.
- Vendor distributions (for example AWS Distro for OpenTelemetry, Grafana Alloy, Splunk, Elastic): curated, supported builds.
- Custom builds with the OpenTelemetry Collector Builder (OCB), which include only the components you need, for a smaller attack surface and binary.
OTTL and Transformation
The OpenTelemetry Transformation Language (used by the transform and filter processors) edits telemetry with statements like set(attributes["tenant"], resource.attributes["k8s.namespace.name"]) or drops health-check spans with IsMatch(attributes["url.path"], "/healthz"). It's how you normalize, enrich, and reduce data without touching application code.
Best Practices
Always Configure memory_limiter and batch
memory_limiter protects the Collector from OOM crashes by refusing data under pressure (clients retry); batch greatly improves export efficiency. They're the two processors every pipeline should have.
Monitor the Collector Itself
Scrape the Collector's internal metrics: accepted, refused, and dropped spans and points, exporter queue size, send failures, and memory. Alert on refused and dropped data and on growing queues, because a silently failing Collector means silently missing telemetry.
Enable Persistent Queues for Critical Paths
Exporters' sending_queue can use the file_storage extension, so buffered data survives restarts and backend outages, trading disk for durability.
Filter and Redact Early
Drop noisy telemetry (health checks, debug logs, unused metrics) and remove sensitive attributes (tokens, emails, card numbers) at the agent or gateway. That reduces cost and compliance risk. See PII handling.
Common Mistakes
Tail Sampling Without Trace-Aware Load Balancing
Spreading one trace's spans across several gateway replicas means each replica sees partial traces and makes inconsistent sampling decisions. Route by trace ID with the loadbalancing exporter.
Using the debug Exporter in Production
The debug exporter (formerly logging) at detailed verbosity prints every span to stdout, which floods logs and burns CPU. Use it only while troubleshooting.
One Giant Pipeline for Everything
Mixing signals from different tenants or sensitivity levels in one pipeline makes redaction, routing, and quotas hard. Split pipelines by purpose, or use routing connectors.
FAQ
Do I need a Collector, or can apps export directly?
Apps can export OTLP directly to many backends, which is fine for small setups. A Collector adds batching, retries, enrichment, sampling, redaction, multi-backend export, and vendor independence without app changes. Most production deployments use one.
What's the difference between agent and gateway mode?
The Collector binary is the same; the deployment differs. Agents run close to workloads (per node or pod) to collect local data and add host context. Gateways run as a central, scalable tier for cross-cutting processing like tail sampling and routing to vendors. Many deployments use both.
Can the Collector replace Prometheus scraping?
The prometheus receiver can scrape Prometheus endpoints using standard scrape configs, and export via remote write or OTLP. It's a good fit when you want one agent for all signals. Prometheus itself still offers the richer local query and alerting engine.
How do I run the Collector on Kubernetes?
Use the OpenTelemetry Operator, which manages OpenTelemetryCollector resources in deployment, DaemonSet, sidecar, or StatefulSet modes, plus Instrumentation resources for automatic SDK injection. The official Helm charts are an alternative.
Related Topics
- OpenTelemetry — The project overview
- OpenTelemetry Sampling — Head and tail sampling strategies
- OpenTelemetry Instrumentation — Producing the data the Collector receives
- Distributed Tracing — Traces end to end
- Prometheus — Exporting metrics to Prometheus-compatible storage
- Log Aggregation — Shipping logs through pipelines