Retries Without a Ceiling: Backoff, Jitter, Dead Letters

Retry layers multiply load on a service that is already struggling. What a ceiling, jitter and a dead-letter queue each do, and what to ask a vendor.

By Harinderpal Hanspal on October 2026

Each layer that retries multiplies the load on a dependency that is already struggling. AWS's worked example, three retries in each of five layers, gives 243 times the load. A ceiling, jittered backoff and a dead-letter queue are the three things that stop it.

Each layer that retries multiplies the load on a struggling service, until a ceiling stops itSketch of one request passing through three boxes in a row, agent, workflow step and tool client, into a slow maintenance system. Each box is marked with three retries, and a red note says the load on the maintenance system is multiplied at every layer. Below, a ceiling line caps the attempts, and an item past the ceiling drops into a box labeled dead letters, where a person looks at it.Agent: 3retriesStep: 3retriesClient: 3retriesSlow systemone request becomes dozens of attemptsceiling: stop after N attemptswait grows, with random jitterDead letters: a person looks
Each layer that retries multiplies the load on a struggling service, until a ceiling stops it

On 7 December 2021 an automated scaling activity inside AWS set off unexpected behavior in a large number of internal clients. AWS's own account says the delays led to "even more connection attempts and retries", and that a latent issue kept those clients from backing off adequately (AWS, post-event summary). Recovery of the affected network devices was complete at 2:22 PM Pacific, after a morning start of 7:30 AM. The retries did not start the fault. They kept it going.

The mechanism is old and the vendors document it plainly. Marc Brooker of AWS wrote that retries are "selfish": a client that retries takes more server time to improve its own odds of success. His example is a five-layer stack of services ending at a database, with three retries in each layer. When the database slows, the load on it rises 243-fold, and recovery becomes unlikely (AWS Builders' Library). Google's SRE book gives a smaller version: three layers, four attempts each, and one user action becomes 64 attempts on the database (Google SRE book).

Where agent workflows add a layer

Picture an agent at a plant that reads open work orders from the maintenance system, drafts a summary and files updates. The maintenance system slows during a shift change. The HTTP client retries. The workflow step retries the client. The engine retries the step. And the agent, reading an error message, decides on its own to call the tool again. Nobody configured that last layer, and nobody counted it.

Microsoft's retry guidance makes the same point about stacks: if a task with a retry policy calls another task with one, the extra layer adds long delays, and the lower-level task should fail fast and report upward (Azure Architecture Center).

What each safeguard does

A ceiling caps attempts. The SRE book suggests limiting retries per request and a budget for the whole process: it gives 60 retries per minute as an example, and past that the request fails instead of retrying. The default matters too. Temporal's documented retry policy has unlimited maximum attempts unless you set one (Temporal). An engine like that hands the ceiling to whoever writes the flow, and a flow written in a hurry has none.

Backoff with jitter spaces the attempts. Exponential backoff widens the gap after each failure, and a cap keeps it from growing absurd. Jitter adds randomness so that a thousand clients that failed together do not all retry together. The SRE book says to "always use randomized exponential backoff". Note that Azure's retry pattern says to spread retries across instances but never uses the word jitter.

A dead-letter queue takes the item that keeps failing out of the loop. In Amazon SQS, a message that cannot be processed moves to a separate queue after a set number of receives, where it can be examined to see why (Amazon SQS). Picture a request to shut a line down that fails five times. It should not retry forever, and it should not vanish. It should land where a person sees it.

A dead-letter queue nobody watches is a graveyard. The same guide covers alarms for messages that arrive there, and warns that a dead-letter queue can break the order of messages when order matters.

Some errors should never retry. Azure's pattern says to cancel when a fault is unlikely to clear. Google's says to return a specific status when overloaded, so callers back off instead of retrying.

Questions to put to a vendor about retries

  1. What is the default maximum number of attempts for a step, and is there one?
  2. Is there a budget across the whole run, or only a limit per step?
  3. Does the backoff include jitter, and is there a cap on the wait?
  4. Which errors are never retried, and who decides?
  5. Where does an item go after its last attempt, and who is alerted?
  6. How many retry layers sit between an agent's request and the system it calls, counting the agent's own?

Drawn from the AWS Builders' Library article on timeouts, retries and backoff with jitter (read on a German-language mirror), Google's SRE book chapter on cascading failures, the Azure retry pattern (18 July 2024), AWS's summary of the 7 December 2021 US-EAST-1 event, the Amazon SQS dead-letter queue guide and Temporal's retry policy page, all read on 6 October 2026. The plant scenario is a hypothetical.

Related notes

Related paper: Governing agents in production: what to ask before an agent acts