Learning · Event-driven & Kafka
Operating Kafka in production
Lag, retention, quotas, and the signals seniors watch so the bus stays boring under load.
Kafka in production is an operations problem as much as a design problem. The log only helps if you can see it.
Watch consumer lag
Lag is the primary health signal for async work. Alert on:
- lag above a business SLO (not just “messages > 0”)
- lag growing while throughput is flat
- partitions stuck while others drain (poison or uneven keys)
Tie lag alerts to user impact: “refunds delayed,” not only “group X behind.”
Retention and disk
Retention is a product decision disguised as a broker setting. Too short and you cannot replay. Too long and you pay for storage and slow rebalances. Match retention to:
- audit / compliance needs
- rebuild windows for projections
- cost
Know whether each topic is delete-retention or compact — and what that implies for history.
Capacity basics
- partition count vs consumer parallelism
- producer throughput and broker ISR health
- consumer fetch sizes and processing time per message
- quotas so one noisy client cannot starve others
Hot keys (one partition absorbing most traffic) show up as lag on a single partition. Fix the key design; more consumers will not help.
Security and access
Treat topics like APIs: authenticate clients, authorize produce/consume per topic, encrypt in transit. Public company clusters with wide-open ACLs are an incident waiting for credentials to leak.
Make it observable
Correlate produce and consume with the same trace / correlation ids. Log event type, key, and outcome. Metrics: produce rate, consume rate, error rate, DLQ rate, end-to-end latency from occurred-at to applied.
A quiet Kafka cluster with clear SLOs is the goal — not a clever topology nobody can operate.