Log Aggregation
Log aggregation is the practice of collecting logs from every server, container, function, and managed service into one searchable place. When a request fails, you want to search a single system for its request ID and see what every service logged about it — not SSH into twelve machines and grep. Aggregated logs also power alerts, security investigations, audits, and dashboards.
The hard parts aren't collecting logs; they're structure, volume, and cost. A busy system can produce terabytes of logs per day, most of which nobody reads. Good log aggregation is as much about deciding what to keep, for how long, and in what form, as it is about tooling.
TL;DR
- Emit structured JSON logs with consistent fields (timestamp, level, service, trace ID, request ID).
- Ship logs with an agent (Fluent Bit, Vector, OpenTelemetry Collector) — not from application code to a remote API.
- Parse, enrich, and redact in the pipeline: add Kubernetes metadata, drop noise, mask PII.
- Pick storage by query needs and cost: Elasticsearch/OpenSearch (full-text indexing), Loki (label index, cheap), or cloud-native services.
- Use retention tiers — hot for days, cheap archive for months — and delete the rest.
- Correlate logs with traces via trace IDs so you can jump between them.
Quick Example
A Fluent Bit configuration on Kubernetes that tails container logs, adds pod metadata, drops health-check noise, and ships to Loki:
And the application side — a structured log line that carries the trace ID:
In Grafana you can query {app="checkout"} | json | level="error" and click the trace ID to open the trace in Tempo.
Core Concepts
The Pipeline
Applications should write to stdout/stderr (or a local file) and let an agent handle delivery, buffering, and retries. That keeps apps simple and avoids losing logs when the backend is slow.
Structured Logging
Free-text lines like User 42 failed login from 1.2.3.4 require fragile regexes to query. Structured logs ({"event":"login_failed","user_id":42,"ip":"1.2.3.4"}) can be filtered and aggregated reliably. See Logging for field conventions and log levels.
Storage Options
Enrichment and Redaction
Pipelines add context the app doesn't know — Kubernetes namespace, pod, node, cloud region, deploy version — and remove what shouldn't be stored: passwords, tokens, card numbers, and personal data. Redacting in the agent means sensitive data never leaves the host. See PII Handling.
Retention and Tiering
Correlation With Traces and Metrics
When every log line includes the trace_id from distributed tracing, you can pivot from a slow trace to its logs and back. The OpenTelemetry SDKs inject trace context into logs automatically in many languages.
Best Practices
Standardize Fields Across Services
Agree on names — service, env, level, trace_id, request_id, user_id — ideally following OpenTelemetry semantic conventions. Consistency makes cross-service queries possible.
Log Events, Not Everything
Log meaningful events and errors with context. Don't log every function entry or every successful health check.
Control Volume at the Source and in the Pipeline
Lower the default level to info in production, sample repetitive debug logs, drop known-noisy lines, and set per-team ingest budgets.
Keep Labels Low-Cardinality (Loki)
Labels like namespace, app, and env are good; user_id or request_id as labels create millions of streams. Keep those in the log body.
Alert on Patterns, Not Single Lines
Alert when error log rates cross a threshold or a specific failure repeats, and prefer metrics for the primary alert. Logs explain; metrics detect.
Separate Security and Audit Logs
Audit trails need tamper resistance and longer retention; ship them to a dedicated, access-controlled store or SIEM.
Common Mistakes
Unstructured, Inconsistent Logs
Every service inventing its own format makes aggregation expensive and queries brittle.
Logging Secrets and Personal Data
Tokens and passwords in logs are a common breach vector, and personal data in logs complicates GDPR deletion requests.
Unbounded Retention
Keeping everything forever in hot storage is how log bills exceed compute bills. Set retention per tier and per log type.
Shipping Directly From Application Code
Synchronous HTTP calls to a log backend slow requests and lose logs during outages. Use a local agent with buffering.
No Timestamps in UTC
Mixed time zones across hosts make incident timelines confusing. Log in UTC with ISO 8601 and millisecond precision.
FAQ
What is log aggregation?
It's collecting logs from all your applications and infrastructure into a central system where they can be searched, correlated, visualized, and alerted on.
What is the ELK stack?
Elasticsearch (storage and search), Logstash (ingestion and processing), and Kibana (visualization). Today it's often "Elastic Stack," with lightweight Beats or Fluent Bit replacing Logstash for collection.
Should I use Loki or Elasticsearch?
Choose Loki when cost and simplicity matter and you mostly filter by service and time before searching text. Choose Elasticsearch or OpenSearch when you need fast full-text search and complex aggregations across very large log sets.
How long should I keep logs?
Keep detailed logs hot for one to two weeks for debugging, cheaper copies for one to three months, and only compliance-relevant logs (audit, security) for longer, as your regulations require.
What's the difference between logs, metrics, and traces?
Metrics are aggregated numbers for detecting problems, traces show a request's path across services, and logs record detailed events for diagnosing specific cases. Aggregated logs are most useful when linked to traces and metrics.
Related Topics
- Logging — What and how to log in application code
- Elasticsearch — A common log storage and search backend
- Grafana — Querying Loki and correlating with traces
- OpenTelemetry — Vendor-neutral log, metric, and trace collection
- Distributed Tracing — Correlating logs with request paths