Learning · Microservices

Consistency and sagas

How to keep business invariants correct across services without distributed transactions — outbox, idempotency, and compensating workflows.

Split data and you split consistency. Senior engineers design for that explicitly instead of hoping the network behaves like one database.

Strong consistency stays local

Inside one service, use local transactions for invariants that must be atomic. Across services, prefer eventual consistency with clear rules for what “caught up” means and what users see while catching up.

Telling product “it will be eventually consistent” is incomplete. Specify: delay expectations, retry behavior, and how to reconcile when something failed halfway.

The dual-write problem

Never write to the database and publish to Kafka (or call a peer) as two independent steps without a plan. One can succeed and the other fail — and you will not know which story is true.

Prefer the transactional outbox (or equivalent): write business state and an outbox record in one transaction; a reliable publisher drains the outbox. Consumers stay idempotent because delivery will retry.

Sagas for multi-step workflows

When a business process spans services (capture → settle → notify → reconcile), model it as a saga: a sequence of local transactions with compensations when a later step fails.

Two common styles:

  • Choreography — services react to each other’s events (simple to start, harder to see the whole flow)
  • Orchestration — one coordinator drives steps (clearer control and visibility; becomes a hotspot if overused)

Pick based on complexity and who owns the process. Payments orchestration often warrants an explicit orchestrator; fan-out notifications often do not.

Idempotency is not optional

At-least-once delivery is the usual reality. Handlers must tolerate duplicates:

  • stable idempotency keys on commands
  • unique constraints / processed-message stores
  • side effects that are safe to replay or gated behind “already done” checks

Without this, retries (your friend in distributed systems) become double charges and duplicate emails.

What to tell the business

Be clear about:

  • which steps are immediate vs eventually reflected
  • how operators reprocess or compensate
  • which failures need human intervention vs automatic retry

Consistency is a product and operations concern, not only an implementation detail.

← Microservices