Log Aggregation

Log aggregation is the practice of collecting logs from every server, container, function, and managed service into one searchable place. When a request fails, you want to search a single system for its request ID and see what every service logged about it — not SSH into twelve machines and grep. Aggregated logs also power alerts, security investigations, audits, and dashboards.

The hard parts aren't collecting logs; they're structure, volume, and cost. A busy system can produce terabytes of logs per day, most of which nobody reads. Good log aggregation is as much about deciding what to keep, for how long, and in what form, as it is about tooling.

TL;DR

Quick Example

A Fluent Bit configuration on Kubernetes that tails container logs, adds pod metadata, drops health-check noise, and ships to Loki:

And the application side — a structured log line that carries the trace ID:

In Grafana you can query {app="checkout"} | json | level="error" and click the trace ID to open the trace in Tempo.

Core Concepts

The Pipeline

Applications should write to stdout/stderr (or a local file) and let an agent handle delivery, buffering, and retries. That keeps apps simple and avoids losing logs when the backend is slow.

Structured Logging

Free-text lines like User 42 failed login from 1.2.3.4 require fragile regexes to query. Structured logs ({"event":"login_failed","user_id":42,"ip":"1.2.3.4"}) can be filtered and aggregated reliably. See Logging for field conventions and log levels.

Storage Options

Enrichment and Redaction

Pipelines add context the app doesn't know — Kubernetes namespace, pod, node, cloud region, deploy version — and remove what shouldn't be stored: passwords, tokens, card numbers, and personal data. Redacting in the agent means sensitive data never leaves the host. See PII Handling.

Retention and Tiering

Correlation With Traces and Metrics

When every log line includes the trace_id from distributed tracing, you can pivot from a slow trace to its logs and back. The OpenTelemetry SDKs inject trace context into logs automatically in many languages.

Best Practices

Standardize Fields Across Services

Agree on names — service, env, level, trace_id, request_id, user_id — ideally following OpenTelemetry semantic conventions. Consistency makes cross-service queries possible.

Log Events, Not Everything

Log meaningful events and errors with context. Don't log every function entry or every successful health check.

Control Volume at the Source and in the Pipeline

Lower the default level to info in production, sample repetitive debug logs, drop known-noisy lines, and set per-team ingest budgets.

Keep Labels Low-Cardinality (Loki)

Labels like namespace, app, and env are good; user_id or request_id as labels create millions of streams. Keep those in the log body.

Alert on Patterns, Not Single Lines

Alert when error log rates cross a threshold or a specific failure repeats, and prefer metrics for the primary alert. Logs explain; metrics detect.

Separate Security and Audit Logs

Audit trails need tamper resistance and longer retention; ship them to a dedicated, access-controlled store or SIEM.

Common Mistakes

Unstructured, Inconsistent Logs

Every service inventing its own format makes aggregation expensive and queries brittle.

Logging Secrets and Personal Data

Tokens and passwords in logs are a common breach vector, and personal data in logs complicates GDPR deletion requests.

Unbounded Retention

Keeping everything forever in hot storage is how log bills exceed compute bills. Set retention per tier and per log type.

Shipping Directly From Application Code

Synchronous HTTP calls to a log backend slow requests and lose logs during outages. Use a local agent with buffering.

No Timestamps in UTC

Mixed time zones across hosts make incident timelines confusing. Log in UTC with ISO 8601 and millisecond precision.

FAQ

What is log aggregation?

It's collecting logs from all your applications and infrastructure into a central system where they can be searched, correlated, visualized, and alerted on.

What is the ELK stack?

Elasticsearch (storage and search), Logstash (ingestion and processing), and Kibana (visualization). Today it's often "Elastic Stack," with lightweight Beats or Fluent Bit replacing Logstash for collection.

Should I use Loki or Elasticsearch?

Choose Loki when cost and simplicity matter and you mostly filter by service and time before searching text. Choose Elasticsearch or OpenSearch when you need fast full-text search and complex aggregations across very large log sets.

How long should I keep logs?

Keep detailed logs hot for one to two weeks for debugging, cheaper copies for one to three months, and only compliance-relevant logs (audit, security) for longer, as your regulations require.

What's the difference between logs, metrics, and traces?

Metrics are aggregated numbers for detecting problems, traces show a request's path across services, and logs record detailed events for diagnosing specific cases. Aggregated logs are most useful when linked to traces and metrics.

Related Topics

References