Orchestration and Choreography Patterns

📖 11 min read

A business process that touches four services has to be driven by something. Orchestration and choreography are the two answers to where that something lives: in one service that calls the others in order, or in the services themselves, each reacting to what the previous one announced.

The saga pattern is a different question layered on top. Orchestration and choreography decide who drives the process. A saga decides what happens when the process gets halfway through and fails, which matters whenever the steps have already changed data in databases no single transaction can reach. A saga can be built either way, so these are two choices, not three.

Orchestration

One service holds the workflow. It calls each participant in turn, holds the state of the process between calls, and decides what happens when a call fails. Participants expose operations and know nothing about the process they are part of.

Order Orchestrator:
  1. Call Payment Service     → wait for response
  2. Call Inventory Service   → wait for response
  3. Call Shipping Service    → wait for response
  4. Call Notification Service
  5. On any failure, run the compensation for every completed step

Use when:

  • The process has enough steps, branches, and retries that nobody can hold the whole thing in their head
  • Someone needs to answer “where is order 4471 right now” without correlating logs
  • Failure handling differs step by step rather than being uniform
  • The sequence itself is a business rule that changes on its own schedule

Example: An order fulfillment process where the orchestrator takes payment, reserves inventory, arranges shipping, and notifies the customer, in that order, and unwinds what it has done if a step fails.

Trade-offs: The orchestrator knows every participant, so it accumulates the coupling the participants shed, and a workflow change usually means changing and redeploying it. It also sits on the path of every process it runs, which makes it both a bottleneck under load and a component whose failure stops work that the participants themselves could have done. Pushed far enough, an orchestrator that holds all the logic and calls services that hold none is a monolith with network calls in the middle.

Dedicated workflow engines exist because the state handling is the hard part rather than the calling. AWS Step Functions, Temporal, and Camunda all persist workflow state, resume after a crash, and handle retries and timeouts, which is most of what a hand-written orchestrator gets wrong.


Choreography

No service owns the process. Each one does its work, publishes an event saying what it did, and other services react to the events they care about. The workflow exists only as the sum of those reactions.

Flow

Orchestration and Choreography

One service calling the others in order, or services reacting to an event.

Orchestration compared with choreography In orchestration, an order orchestrator holds the process state and calls payment, inventory, shipping, and notification services in order, waiting for each response. The participants know nothing about the process. In choreography, the order service publishes an OrderCreated event to an event bus, and the inventory, payment, shipping, and notification services each react to it. No service owns the process. ORCHESTRATION Order orchestrator holds the process state Payment 1 Inventory 2 Shipping 3 Notification 4 Calls each in order and waits; participants know nothing about the process. CHOREOGRAPHY Order service event bus OrderCreated Payment Inventory Shipping Notification Each reacts on its own; the workflow is only the sum of those reactions. call and response event

Use when:

  • Steps are genuinely independent and don’t need to happen in a fixed order
  • New participants should be able to join by subscribing, without a change anywhere else
  • Services are owned by teams that shouldn’t have to coordinate a release to add a reaction
  • The process is short enough that no one needs a central view of it

Example: Placing an order publishes OrderCreated, and the inventory, payment, shipping, and notification services each pick it up and act without anything telling them to.

Trade-offs: The process exists but is written down nowhere, so understanding it means reading every subscriber, and answering what happened to one order means correlating events across services. Adding a subscriber is easy in exactly the way that makes cycles easy to create, where service A’s event triggers B, whose event triggers A. Failure handling is also distributed, so each participant has to decide for itself what to do about a step it didn’t perform and can’t see.


Choosing Between Them

  Orchestration Choreography
Where the workflow is written In one place, as explicit steps Nowhere, so it emerges from what each service subscribes to
Adding a step Change and redeploy the orchestrator Deploy a new subscriber, often touching nothing else
Finding out what happened Read the orchestrator’s state for that instance Correlate events across services, which needs tracing to be bearable
Who knows about whom The orchestrator knows every participant; participants know nobody Nobody knows anybody, but everybody knows the event shapes
How it fails badly Becomes a bottleneck, or a monolith with the services as libraries Becomes an undocumented workflow with cycles nobody designed

The split is not usually all-or-nothing. A common arrangement orchestrates the part of a process that has a required order and real compensation, and lets everything downstream of it, such as notifications, analytics, and search indexing, happen by choreography.


Saga

Introduced by Hector Garcia-Molina and Kenneth Salem (1987), and adapted for microservices by Chris Richardson among others

A saga is a sequence of local transactions, each committed in one service’s own database, with a compensating action defined for each one. If a later step fails, the saga runs the compensations for the steps that already committed, in reverse order.

The problem it solves: In a single database, one transaction covers every write and either all of it happens or none does. Across services with separate databases, there is no such transaction. Two-phase commit exists, but it holds locks across every participant for the whole duration and stalls if the coordinator dies mid-commit, and the datastores services actually use, including document stores, message brokers, and many managed cloud services, frequently don’t offer it at all.

Structure

One Transaction Versus Separate Databases

One all-or-nothing commit, or three local commits with nothing tying them.

A single database transaction compared with separate service databases In a monolith, one transaction inserts the order, updates inventory, and inserts the payment, and commits all or nothing. With microservices, the order, payment, and inventory services each own a database and commit locally, and no shared transaction covers the three. ONE DATABASE BEGIN TRANSACTION INSERT order UPDATE inventory INSERT payment COMMIT: all or nothing SEPARATE DATABASES Order own DB local commit Payment own DB local commit Inventory own DB local commit No transaction covers all three. transaction boundary

How a saga runs:

Flow

A Saga Succeeding and Compensating

Each step commits locally; a failure runs compensations in reverse.

Saga happy path and failure path with compensation On the happy path, the saga creates the order as pending, reserves inventory, charges payment, and completes the order. On the failure path, the payment charge is declined at step 3, so the saga runs the compensations for the steps that already committed in reverse order: first releasing inventory, the compensation for step 2, then cancelling the order, the compensation for step 1. HAPPY PATH: EVERY STEP COMMITS Create order T1, pending Reserve inventory T2 Charge payment T3 Complete order confirmed STEP 3 FAILS: COMPENSATE IN REVERSE Create order T1 Reserve inventory T2 Charge payment T3 ✗ payment declined Release inventory C2, undoes T2 Cancel order C1, undoes T1 first then

Compensating Transactions

A compensation is not a rollback. The original transaction has committed and other work has happened since, so the compensation is a new transaction that makes business sense of the reversal rather than pretending the first one never ran.

Original transaction Compensation Why it isn’t a rollback
Create order Cancel order The order id is already issued and may have been shown to the customer, so it is marked cancelled, not deleted
Reserve inventory Release inventory Other orders have reserved and released stock since, so the compensation returns quantity rather than restoring a prior state
Charge payment Refund payment A settled charge can’t be withdrawn, so a refund is a separate movement of money that both parties can see
Send email Send correction Nothing can unsend it

Compensatable, Pivot, and Retriable

Not every step can be undone, which means a saga has a point of no return. Richardson’s taxonomy names the three kinds of step, and identifying the pivot is the part of saga design that is easy to skip and expensive to get wrong.

  • Compensatable transactions run before the point of no return and each have a compensation that can undo them.
  • The pivot transaction is the go/no-go point. Once it commits, the saga is committed to finishing, so there is exactly one of these. It may be the last compensatable step, the first retriable one, or a step that is neither.
  • Retriable transactions come after the pivot. They cannot be undone, so the saga has to keep retrying each one until it succeeds, which means they must be designed so that succeeding is always eventually possible.
Flow

Compensatable, Pivot, and Retriable Steps

The pivot is the point of no return in a saga.

Compensatable, pivot, and retriable saga steps Five saga steps in order. CreateOrder and ReserveInventory are compensatable, with CancelOrder and ReleaseInventory as their compensations. ChargePayment is the pivot, the go or no-go point. After it commits, the saga is committed to finishing. ShipOrder and SendConfirmation are retriable: no compensation exists, so the saga retries each until it succeeds. T1 CreateOrder C1: CancelOrder T2 ReserveInventory C2: ReleaseInventory T3 ChargePayment C3: RefundPayment T4 ShipOrder no compensation T5 SendConfirmation no compensation COMPENSATABLE: CAN BE UNDONE PIVOT RETRIABLE: RETRIED UNTIL THEY SUCCEED point of no return: once the pivot commits, the saga is committed to finishing

Placing the pivot is a business decision rather than a technical one. Charging before shipping makes the charge the last reversible step, and everything after it has to be something the business is willing to retry until it works.

Orchestrated and Choreographed Sagas

Both coordination styles from earlier in this guide apply to sagas, and the comparison table above holds here too. Compensation raises what is at stake, because failure handling is exactly what the two styles place differently.

Orchestrated: the orchestrator holds the saga state and calls compensations itself when a step fails.

C4 · Dynamic

An Orchestrated Saga

The orchestrator holds saga state and calls the compensations itself.

Orchestrated saga with compensation A saga orchestrator holds the saga state for order 123. It calls the order service to create the order, which succeeds, then the inventory service to reserve stock, which succeeds, then the payment service to charge, which fails. It then calls the inventory service to release the stock, compensating step 2, and the order service to cancel, compensating step 1. The shipping service is never called. Saga orchestrator saga state: orderId 123, step PAYMENT Order service Inventory service Payment service Shipping service 1 create 5 cancel 2 reserve 4 release 3 charge: FAILED never called 4 compensates T2 and 5 compensates T1, in reverse order. forward step failed step compensation

Choreographed: each service reacts to events, and a failure event is what triggers the compensations upstream of it.

Flow

A Choreographed Saga

A failure event travels upstream and each service compensates itself.

Choreographed saga with compensation The order service publishes OrderCreated, the inventory service reserves stock and publishes InventoryReserved, and the payment service fails and publishes PaymentFailed. That failure event reaches every upstream participant: the inventory service releases stock and publishes InventoryReleased, and the order service cancels and publishes OrderCancelled. No single place holds the saga state. Order service Inventory service Payment service OrderCreated InventoryReserved PaymentFailed cancels, then publishes OrderCancelled releases, then publishes InventoryReleased Saga state is implied by which events have been published. No single place holds it. event failure event, sent upstream

Each service listens for the events it cares about, commits its local transaction, and publishes the result.

The saga state in the choreographed version is implied by which events have been published and which have not, so there is no single place to query how far a given order has progressed. That is the cost people underestimate.

When a Compensation Fails

Nothing compensates a compensation, so a failed compensation leaves the system in a state the saga cannot resolve on its own.

T1 ✓ → T2 ✓ → T3 ✗ → C2 ✗   (compensation itself failed)

Compensations therefore have to be retriable, which means they have to be idempotent, because a retry can’t tell whether the previous attempt partly succeeded. ReleaseInventory(orderId) called twice must release the stock once. Where retries are exhausted, the remaining options are to escalate to a human with enough context to resolve it by hand, or to complete the saga forward if finishing is less damaging than staying half-done.

Saga Limitations

No isolation. A saga's intermediate states are visible to everyone else, so another reader can see an order that exists with a payment that hasn't happened. Semantic locks, such as an explicit pending status that other operations check, are the usual answer.

Every step doubles. N steps means N compensations to write, test, and keep correct as the business logic they reverse changes.

The window is visible to everyone. The system is inconsistent for as long as the saga runs, which can be seconds or, where a step waits on a human or an external provider, considerably longer.


Quick Reference

  What it decides Reach for it when Main cost
Orchestration Where the workflow logic lives The sequence is complex, or someone must be able to see process state A component every process runs through, holding all the coupling
Choreography Where the workflow logic lives Participants are independent and should be addable without coordination A workflow that exists nowhere and can only be reconstructed
Saga What happens when a multi-service process fails partway Steps commit to separate databases and a partial result is unacceptable A compensation per step, no isolation, and a visible inconsistency window

Found this guide helpful? Share it with your team:

Share on LinkedIn