Learning · Event-driven & Kafka
Failure, DLQ, and replay
Retries, dead letters, and rebuilding state from the log — recovery without tribal knowledge.
Event-driven systems fail in public. Plan for poison messages, stuck consumers, and “we need yesterday’s projection again.”
Separate transient from poison
Transient: broker blips, DB timeouts, downstream 503s → retry with backoff, then resume.
Poison: bad schema, invariant violation, bug that always throws → do not block the partition forever. After a budgeted retry count, route to a dead-letter topic (or equivalent) and continue.
Dead letters need owners
A DLQ without alerts is a trash can. Track:
- depth and age
- error reason
- original topic / partition / offset
- correlation id
Someone must be on-call to classify, fix, and replay — or deliberately discard with a recorded decision.
Replay is a feature
Kafka’s value is the log. Use it:
- to rebuild a projection into a new store
- to onboard a new consumer from earliest or a chosen timestamp
- to reprocess after a bugfix (with idempotent applies)
Document retention so replay stays possible for the window you care about. Compacted topics are not a full audit history unless you designed them that way.
Backpressure and lag
When consumers fall behind, decide consciously: scale out (if partitions allow), shed non-critical work, or pause producers. Blindly growing lag until the disk fills is not a strategy.
Runbooks over heroics
Write down: how to pause a consumer group, how to reset offsets safely, how to replay a DLQ batch, who owns each topic. Incident time is too late to invent that.