Learning · Event-driven & Kafka

Failure, DLQ, and replay

Retries, dead letters, and rebuilding state from the log — recovery without tribal knowledge.

Event-driven systems fail in public. Plan for poison messages, stuck consumers, and “we need yesterday’s projection again.”

Separate transient from poison

Transient: broker blips, DB timeouts, downstream 503s → retry with backoff, then resume.

Poison: bad schema, invariant violation, bug that always throws → do not block the partition forever. After a budgeted retry count, route to a dead-letter topic (or equivalent) and continue.

Dead letters need owners

A DLQ without alerts is a trash can. Track:

  • depth and age
  • error reason
  • original topic / partition / offset
  • correlation id

Someone must be on-call to classify, fix, and replay — or deliberately discard with a recorded decision.

Replay is a feature

Kafka’s value is the log. Use it:

  • to rebuild a projection into a new store
  • to onboard a new consumer from earliest or a chosen timestamp
  • to reprocess after a bugfix (with idempotent applies)

Document retention so replay stays possible for the window you care about. Compacted topics are not a full audit history unless you designed them that way.

Backpressure and lag

When consumers fall behind, decide consciously: scale out (if partitions allow), shed non-critical work, or pause producers. Blindly growing lag until the disk fills is not a strategy.

Runbooks over heroics

Write down: how to pause a consumer group, how to reset offsets safely, how to replay a DLQ batch, who owns each topic. Incident time is too late to invent that.

← Event-driven & Kafka