Event Sourcing
Most applications store current state: a row in an accounts table says the balance is $420. How it got there (deposits, withdrawals, corrections) is lost, unless you add a separate audit log that may drift out of sync. Event sourcing flips this: the source of truth is an append-only sequence of events (AccountOpened, MoneyDeposited, MoneyWithdrawn), and current state is derived by replaying them.
That gives you a perfect audit trail, the ability to reconstruct state at any point in time, and the freedom to build new read models from history. It also brings real complexity: event versioning, eventual consistency in read models, and a different way of thinking about data. Event sourcing pairs naturally with CQRS and domain-driven design, and fits best in domains where history and business intent genuinely matter.
TL;DR
- State = fold over events. The event log is the source of truth, and it's append-only and immutable.
- Events are grouped into streams, usually one per aggregate instance (such as
account-42). - Commands are validated against current state, then produce new events, appended with optimistic concurrency (the expected version).
- Projections build query-optimized read models from events, which are eventually consistent. That's CQRS.
- Snapshots speed up rebuilding long streams. Upcasting handles old event versions.
- Worth it for audit-heavy, temporal, or complex domains. Overkill for simple CRUD.
Quick Example
A minimal event-sourced aggregate in TypeScript:
Core Concepts
Events as the Source of Truth
A domain event records a fact that happened, named in the past tense with business meaning: OrderPlaced, ShipmentDispatched, PriceAdjusted. Events are immutable. Mistakes are corrected with new compensating events (DepositReversed), never by editing history. See domain events.
Streams and Aggregates
Events are stored in streams, typically one per aggregate instance: the consistency boundary from DDD. Loading an aggregate means reading its stream and applying events in order (evolve). Commands are handled by decide(state, command) → events. This decide/evolve split keeps the business logic pure and very testable: given past events, when a command, then expect new events.
Optimistic Concurrency
When appending, the writer supplies the stream version it read. If another writer appended in the meantime, the append fails with a conflict, and the handler reloads and retries (or reports a conflict). That enforces invariants per aggregate without locks.
Projections and Read Models
Rebuilding state from events is fine for one aggregate, but not for queries like "all overdue invoices". Projections subscribe to the event log and maintain read models (SQL tables, search indexes, caches) optimized for queries. They're eventually consistent: typically milliseconds to seconds behind. Because events are retained, you can rebuild projections from scratch, or create entirely new ones from full history. That's CQRS in practice.
Projections must be idempotent and track their position (a checkpoint), so they can resume after restarts. See idempotency.
Snapshots
Aggregates with thousands of events take longer to load. Snapshots periodically store the derived state at a version, so loading reads the latest snapshot plus subsequent events. Treat snapshots as a cache: they can be discarded and rebuilt. Short-lived aggregates rarely need them. A long stream can also signal that the aggregate boundary should be redesigned (for example, closing the books per period).
Event Stores
A common architecture: the event store is the system of record, and events are published to Kafka for other services, via the outbox pattern or store subscriptions.
Versioning Events
Event schemas evolve, but old events live forever. Strategies: add optional fields (backward compatible), upcasters that transform old event versions into new shapes when reading, new event types for new meanings, and rarely, copy-and-transform migrations of streams. See event schema evolution.
Personal Data and Deletion
Immutable logs collide with "right to be forgotten" requirements. Techniques: keep personal data out of events (store references to a deletable store), or crypto-shredding: encrypt personal fields with per-subject keys, and delete the key to render the data unreadable.
When to Use Event Sourcing
Good fits: financial ledgers, accounting, and payments; order and fulfillment lifecycles; compliance-heavy domains requiring audit; collaborative or workflow systems where history matters; domains needing temporal queries ("what did we know on March 3rd?").
Poor fits: simple CRUD apps, systems where history has no business value, and teams without experience in eventual consistency. You can event-source one bounded context without adopting it everywhere.
Best Practices
Model Events on Business Language
InvoiceSent and PaymentReceived capture intent. InvoiceUpdated with a diff of fields loses meaning, and turns the log into a change-data dump.
Keep Aggregates Small
Small aggregates mean short streams, fewer concurrency conflicts, and clear invariants. Cross-aggregate processes are coordinated with events and sagas, not large transactions.
Test With Given/When/Then
Event-sourced logic is ideal for behavior tests: a list of prior events, a command, and the expected new events or error. Fast, deterministic, and readable to domain experts.
Plan Projection Rebuilds
Make projections rebuildable by design, and practice rebuilding them, including blue-green swaps of read models for large histories.
Common Mistakes
Event Sourcing Everything
Applying it to every CRUD entity adds complexity without benefit. Use it where history and domain behavior justify the cost.
Querying the Event Store for Reads
Scanning events to answer list and search queries doesn't scale. Build projections for queries.
Leaking Internal Events as Public Contracts
Fine-grained internal events change often. Publish stable integration events to other services, separate from internal event-sourcing events, so you can refactor internally without breaking consumers.
FAQ
What is the difference between event sourcing and event-driven architecture?
Event-driven architecture is about services communicating by publishing and reacting to events. Event sourcing is a persistence pattern where a service stores its own state as a sequence of events. You can have either without the other, though they complement each other well.
Is event sourcing the same as CQRS?
No. CQRS separates write and read models. Event sourcing stores state as events. They're often combined, because an event log naturally feeds projections into read models, but CQRS works with ordinary databases, and event sourcing can technically be used without separate read models.
Can I use Kafka as an event store?
Kafka is excellent for distributing events, but it lacks per-aggregate optimistic concurrency and efficient reads of an individual entity's stream, so it's usually a poor primary event store. Many systems use a database or dedicated event store for the source of truth, and Kafka for integration.
How do I fix a wrong event?
You don't edit it. Append a compensating or correcting event that reverses or adjusts its effect (PaymentRefunded, AddressCorrected). History stays accurate about what happened, including the correction.
Related Topics
- Event-Driven Architecture — Pillar overview
- CQRS — Separate read and write models
- Domain Events — Designing meaningful events
- Event Schema Evolution — Versioning events that live forever
- Outbox Pattern — Publishing events reliably
- Domain-Driven Design — Aggregates and bounded contexts