Picture booking a holiday that needs a flight, a hotel, and a rental car — three separate companies, any of which might be slow, fail, or quietly drop your request. A good travel agent doesn't just fire off three bookings and hope. They track each one, chase the ones that don't confirm, and if the hotel falls through after the flight is booked, they sort out a fix rather than leaving you stranded.
The Scheduler Agent Supervisor pattern builds that diligent travel agent into your system. It coordinates a sequence of distributed steps, keeps durable track of how each one is going, and — crucially — has a dedicated watcher whose job is to notice when a step has gone wrong and put things right.
The problem
A single business action often spans several remote services, and any of them can fail in messy ways: time out, return an error, succeed but never send back a confirmation, or just hang. In a distributed system these aren't edge cases — they're routine. If you simply call each step in turn with no memory of progress, a crash halfway through leaves the work in an unknown, half-finished state with no one responsible for cleaning it up.
Worse, transient failures and permanent failures look the same in the moment. A naive flow either gives up too early on a blip or retries forever on something that will never succeed. What's missing is something that remembers the intended outcome of every step, compares it to reality, and actively shepherds stalled work toward completion or a clean rollback.
Step through a booking app that keeps its progress in memory, and predict what survives when it crashes mid-trip.
How it works
The pattern splits the work across three roles. The scheduler starts the workflow, breaks it into steps, and writes the whole plan plus each step's status to a durable state store, so progress survives a crash. Each step is carried out by an agent: an isolated worker that talks to one remote service and reports back, shielding the rest of the system from that service's quirks.
When the scheduler hands a step to an agent, it records a lease: a deadline by which the agent must report back or renew it. The lease is what makes a silent failure visible. A crashed agent can't tell anyone it crashed, but its lease still runs out.
The star is the supervisor. It scans the state store on a timer, looking for steps that are stuck, failed, or past their lease. For a transient failure it puts the step back to be retried, with a limit on attempts. For one that will never succeed, it triggers compensation to undo the steps already done and leave the system consistent. Because every status is recorded, the supervisor can pick up after any interruption — even its own crash.
Step through a trip booking below. An agent dies halfway through step 2; predict what the supervisor does when the lease expires. At the end, switch how the hotel API behaves on the retry to see both endings: a successful retry, or compensation.
Make every step idempotent and re-runnable. The supervisor's recovery only works safely if retrying a step that may have partly succeeded doesn't double-charge or duplicate. Design agents so that running the same step twice lands you in the same place as running it once.
An agent crashed halfway through charging a card, and the supervisor retried the step. The customer was charged twice. What was missing?
An expired lease doesn't mean the agent stopped. A slow agent, or one cut off by a network blip, can finish its step after the supervisor has handed it to someone else — and then two agents report on the same step. Give each lease an attempt number and have the state store reject reports from an older attempt, so a late finisher can't overwrite the retry's result.
When to use it
Reach for this pattern when a workflow spans multiple unreliable remote services and you need it to either finish completely or unwind cleanly — order fulfilment, provisioning across cloud resources, multi-party financial transactions. It's closely related to the saga: both coordinate distributed steps and lean on compensating transactions to roll back. The distinguishing feature here is the explicit supervisor — a built-in watchdog that turns ordinary retry and recovery into a continuous, self-healing process rather than something you bolt on per call.
The cost is real complexity: a durable state store, a separate supervisor process, and the discipline of idempotent steps. For a quick local transaction or a workflow where partial failure is harmless, that machinery is overkill. But for long-running, high-stakes processes that absolutely must not be left half-done, the supervisor's tireless watching is exactly what keeps the system honest.
A provisioning workflow created a VM and a database, then failed for good at step 3, configuring DNS. What does it mean for the supervisor to compensate?