Learning · Event-driven & Kafka

Operating Kafka in production

Lag, retention, quotas, and the signals seniors watch so the bus stays boring under load.

Kafka in production is an operations problem as much as a design problem. The log only helps if you can see it.

Watch consumer lag

Lag is the primary health signal for async work. Alert on:

  • lag above a business SLO (not just “messages > 0”)
  • lag growing while throughput is flat
  • partitions stuck while others drain (poison or uneven keys)

Tie lag alerts to user impact: “refunds delayed,” not only “group X behind.”

Retention and disk

Retention is a product decision disguised as a broker setting. Too short and you cannot replay. Too long and you pay for storage and slow rebalances. Match retention to:

  • audit / compliance needs
  • rebuild windows for projections
  • cost

Know whether each topic is delete-retention or compact — and what that implies for history.

Capacity basics

  • partition count vs consumer parallelism
  • producer throughput and broker ISR health
  • consumer fetch sizes and processing time per message
  • quotas so one noisy client cannot starve others

Hot keys (one partition absorbing most traffic) show up as lag on a single partition. Fix the key design; more consumers will not help.

Security and access

Treat topics like APIs: authenticate clients, authorize produce/consume per topic, encrypt in transit. Public company clusters with wide-open ACLs are an incident waiting for credentials to leak.

Make it observable

Correlate produce and consume with the same trace / correlation ids. Log event type, key, and outcome. Metrics: produce rate, consume rate, error rate, DLQ rate, end-to-end latency from occurred-at to applied.

A quiet Kafka cluster with clear SLOs is the goal — not a clever topology nobody can operate.

← Event-driven & Kafka