Event Schema Evolution
In an event-driven system, events are contracts between teams that often don't coordinate deployments: a producer publishes OrderPlaced, and a dozen consumers (billing, analytics, email, search) read it. Events are also durable: they sit in Kafka topics for days or forever, and in event-sourced systems they live as permanent history. When the producer changes the event's shape, every consumer, and every stored event, must keep working.
Schema evolution is the discipline of changing event formats safely. It borrows ideas from API versioning, but it's stricter: consumers may read old and new events side by side, you can't "force upgrade" data already written, and there's no request-response negotiation.
TL;DR
- Define events with explicit schemas (Avro, Protobuf, or JSON Schema), and register them in a schema registry.
- Backward compatible: new consumers can read old events. Forward compatible: old consumers can read new events. Full: both.
- Safe changes: add optional fields with defaults, and deprecate (don't remove) fields. Breaking: rename, change types, remove required fields, change meaning.
- For breaking changes, introduce a new event type or version, publish both during migration, or use a new topic.
- For stored events, use upcasters to translate old versions on read.
- Enforce compatibility in CI and at produce time, and use consumer contract tests.
Quick Example
An Avro schema evolved in a backward- and forward-compatible way:
Compatibility enforced by the registry, and checked in CI:
Core Concepts
Compatibility Modes
For shared event streams with independent teams and long retention, full transitive compatibility is the safest default.
Safe and Breaking Changes
Format-Specific Rules
- Avro: the reader and writer schemas are resolved at read time. Fields need defaults to be added or removed compatibly, and aliases support renames. Compact binary, with the schema ID in the message header.
- Protobuf: fields are identified by number, not name. Never reuse or change field numbers, mark removed fields
reserved, and treat all fields as optional (proto3). Unknown fields are preserved by default. It's the same discipline as gRPC APIs. - JSON Schema: flexible, but compatibility depends on
additionalPropertiesand required-field choices. Tolerant readers (ignore unknown fields) are essential.
Schema Registries
A schema registry (Confluent Schema Registry, Apicurio, AWS Glue Schema Registry) stores versioned schemas per subject (usually per topic value), enforces compatibility on registration, and lets serializers embed a small schema ID in each message. Consumers fetch schemas by ID to deserialize. Combined with producer settings that forbid auto-registration in production, it stops incompatible changes before they reach the topic.
Handling Breaking Changes
When a change can't be made compatible:
- New event type or version:
OrderPlacedV2(orcom.acme.orders.v2.OrderPlaced), published alongside v1 during a migration window. Consumers migrate, then v1 is retired. - New topic:
orders.v2, with a bridge that translates v1 ↔ v2 while consumers move. - Expand and contract: add the new field and populate both, migrate consumers to the new field, then stop populating and remove the old one (after retention expires).
Communicate deprecation timelines, and track which consumers still read old versions.
Upcasting Stored Events
In event-sourced systems, old events stay in the store forever. Upcasters transform old versions into the current shape when events are loaded (OrderPlaced v1 → v2: split name into first_name/last_name, and default currency), so domain code only handles the latest version. Keep upcasters small, tested, and chained (v1 → v2 → v3).
Envelopes and Metadata
Standardize an envelope with metadata separate from the payload: event ID, type, schema version, source, timestamp, correlation and causation IDs, and the partition key. CloudEvents is a vendor-neutral specification for this metadata, with bindings for Kafka, HTTP, and other transports. See domain events.
Best Practices
Treat Events as Public APIs
Review schema changes like API changes: owners, changelogs, deprecation policies, and documentation. The producer team owns the contract, and consumers depend on it.
Be a Tolerant Reader
Consumers should ignore unknown fields, handle unknown enum values gracefully, and read only the fields they need. Tolerant consumers make forward-compatible evolution possible.
Test Compatibility in CI
Run registry compatibility checks (or tools like Buf's breaking for Protobuf) on every pull request that changes a schema. Add consumer-driven contract tests so consumers declare the fields they rely on.
Never Change Meaning Silently
Changing units, semantics, or the interpretation of a field is the most dangerous change, because schemas can't detect it. Add a new field with a new name instead.
Common Mistakes
Schemaless JSON Events
Without schemas, every change is a guess about what consumers parse, and breakage shows up in production. Even with JSON, publish JSON Schemas and validate at produce time.
Checking Compatibility Only Against the Latest Version
With long retention or replays, consumers may read events from several versions back. Use transitive compatibility.
Reusing Protobuf Field Numbers
Reusing a removed field's number makes old data deserialize into the wrong field. Mark removed numbers and names as reserved.
FAQ
What is the difference between backward and forward compatibility?
Backward compatibility means new readers can read old data, so you upgrade consumers first. Forward compatibility means old readers can read new data, so you upgrade producers first. Full compatibility supports both, which lets producers and consumers deploy in any order.
Should I use Avro, Protobuf, or JSON Schema for events?
Avro is popular in the Kafka ecosystem, for compact encoding and rich schema resolution. Protobuf suits teams already using gRPC, and it offers strong tooling (Buf) and code generation. JSON Schema is human-readable and easy to adopt, but larger and looser. Any works well with a schema registry and discipline.
How do I rename a field in an event?
Don't rename it in place. Add the new field, populate both during a migration period, move consumers to the new field, then deprecate and later remove the old one. Avro aliases can help readers map old names, but coordinate carefully.
Do I need a schema registry?
For small systems with few consumers, schemas in a shared repository with CI checks can suffice. As the number of teams, topics, and consumers grows, a registry that enforces compatibility at produce time prevents incompatible events from ever being published.
Related Topics
- Event-Driven Architecture — Pillar overview
- Domain Events — Designing event contents
- Event Sourcing — Events that live forever
- API Versioning — The request-response equivalent
- Kafka — Topics, retention, and serializers
- gRPC — Protobuf evolution rules