Workflow Restart Without Repeating a Step: What to Ask
A restart can issue a work order twice or start a shutdown again. What durable engines record, why a record alone is not enough, and what to ask a vendor.
By Harinderpal Hanspal on October 2026
A workflow engine that records each finished step can resume after a restart, but the step in flight when it died may run again. Engine documentation says as much. Safety comes from an idempotency key the receiving system honors, plus timers that live outside the process.
Picture a maintenance workflow at 02:10 on a night shift. It has just issued a work order for a failing pump and is about to book a technician. The server hosting it restarts. When it comes back, does it issue the work order again?
The vendors of workflow engines say it can. Temporal's documentation warns that Activities "may be executed more than once" (Temporal, Activity Definition). AWS describes its Express Workflows as at-least-once, so an execution "could potentially run more than once" (AWS Step Functions). Stripe's engineering team gave the cost of the plain version in 2017: an endpoint that charges a customer, called twice, leaves the customer double-charged (Stripe, 22 February 2017). Swap the charge for a work order, a valve command or a line changeover and the arithmetic is the same.
What a restart has to recover
A script keeps its place in memory, and a restart erases it. A durable engine writes each step to a log as it happens and rebuilds its position from that log. Temporal calls its version Event History, an append-only log that lets a run "recover from a crash and continue making progress" (Temporal). Restate calls it a journal and says it replays it, skipping completed steps (Restate).
Replay has a price. Microsoft's Durable Functions documentation says the orchestrator code runs again and again, so it must produce the same result each time, which rules out reading the clock or generating a random number directly (Microsoft Learn, updated 24 August 2026). The same page warns that a code change can break replay for runs already in progress. Picture deploying new changeover logic while a three-day run is mid-flight.
The step that was in flight
Replay skips steps whose results are on the record. The step that was running when the power went has no result on the record. The log says it started. It cannot say whether the maintenance system accepted the work order, so the engine runs it again.
That is why Step Functions' own page limits its exactly-once model. Standard Workflows never run a task twice unless the flow has retries configured, and the page says the at-least-once Express type suits idempotent actions.
The standard fix is an idempotency key: a unique ID the caller sends with the request, so the receiver can tell a repeat from a new order. Stripe's post describes the pattern, and Temporal suggests building the key from the workflow run ID and the activity ID, which stay the same across retries. Restate says duplicates carrying the same key "return the same result as the original request."
A key works only if the system on the other end honors it. Many maintenance systems and plant historians do not. Then the workflow has to look before it writes: search for an open order carrying its own reference, and create one only if none exists. Microsoft's retry guidance describes the failure this prevents: a service processes a request, fails to send the response, and the retry logic sends it again (Azure Architecture Center).
Timers that outlive the process
Plant workflows wait. Hold the changeover until the line is empty. Wait 72 hours for a part. A sleep call in ordinary code dies with the process. Microsoft tells developers to use durable timers instead, and describes a timer as a message that becomes visible only at its due time, so an app scaled down to zero is woken then (Microsoft Learn). Temporal says its timers are persisted: if the worker or service is down when one comes due, the sleep resolves once both are back (Temporal).
Two engines, two designs, one question left open: what a timer that came due during the outage does when the system returns, and whether the step after it is safe to run late.
Questions to put to a vendor about restarts
- Where is the step history stored, and does it survive the loss of the machine running the workflow?
- For a step that was in flight at the restart, does the engine run it again, and can you see which ones?
- Does every outbound write carry an idempotency key, and does the receiving system honor it?
- Where the receiver has no key support, what check runs before the write?
- What happens to a timer that came due while the system was down?
- If you deploy new logic while a run is mid-flight, what happens to that run?
Drawn from the Temporal documentation (activity definition, event history, timers, read 6 October 2026), Restate's key concepts page, AWS Step Functions' workflow-type page, Microsoft Learn's Durable Functions pages (updated 24 August and 22 April 2026), the Azure retry pattern (18 July 2024) and Stripe's idempotency post (22 February 2017), all read on 6 October 2026. The plant scenarios are hypotheticals.
Related notes
- Where to put the human in an AI agent's work: approve the write, not the draft
- An AI agent log can prove nothing changed and still not say who acted
Related paper: Governing agents in production: what to ask before an agent acts