← Back to the library
System designDeep dive · 9 min

A brief blip turns into a long outage

“A dependency hiccups for seconds, but traffic to it multiplies, error rates stay high after the root cause is gone, and it only recovers when load is shed manually.”

ERROR CODES YOU MAY SEE

A code is a clue. Use its meaning and surrounding evidence to narrow the cause.

503HTTP status codes · RFC 9110 · standard

The server is temporarily unable to handle the request due to overload or maintenance. It may send Retry-After.

Treat it as transient: honor Retry-After when present, otherwise back off with jitter and cap total attempts.

Official reference for 503 (opens in new tab)Code definition reviewed
DEADLINE_EXCEEDEDgRPC status codes · standard

The deadline expired before the operation could complete (status 4). For state-changing operations it may be returned even if the operation completed.

Check whether the deadline was spent waiting in queues or on retries downstream. Propagate the caller's remaining deadline instead of starting a fresh timeout at each hop, and only retry idempotent calls.

Official reference for DEADLINE_EXCEEDED (opens in new tab)Code definition reviewed

THE PRINCIPLE

A retry is extra load. Every retry policy needs a cap, a budget, jitter, and a deadline it cannot outlive.

FIRST MOVES

  1. Map who retries: browser, SDK, gateway, service mesh, each service, the DB driver. Three layers of 3 retries is up to 64 attempts per user action.
  2. Retry at one layer only (usually closest to the caller that owns the intent); set the others to zero.
  3. Use capped exponential backoff with jitter, and a retry budget/token bucket so retries stop when most calls fail.
  4. Propagate one absolute deadline downstream; drop work whose caller has already given up.
  5. Shed load early and cheaply (429/503 with Retry-After) and add circuit breakers so callers fail fast instead of queueing.
  6. Retry only idempotent operations or ones carrying an idempotency key.

TOOLS: THEN → NOW

Per-request retry counts alonePlus a shared retry budget: AWS SDK retry quota, gRPC retryThrottling, Envoy retry_budget
Independent timeouts per hopPropagated deadlines (gRPC deadlines, context/AbortSignal)
Callers wait out every timeoutCircuit breakers in the client library or mesh (Resilience4j, Polly, Envoy)

PATTERN SNAPSHOT

// The only retry layer: capped attempts, full jitter, shared deadline.
async function call<T>(fn: (s: AbortSignal) => Promise<T>, deadline: number) {
  for (let attempt = 0; ; attempt++) {
    try {
      return await fn(AbortSignal.timeout(Math.max(0, deadline - Date.now())));
    } catch (err) {
      if (attempt >= 2 || !isRetryable(err) || !retryBudget.tryAcquire()) throw err;
      const backoff = Math.random() * Math.min(2_000, 100 * 2 ** attempt);
      if (Date.now() + backoff >= deadline) throw err;
      await new Promise((r) => setTimeout(r, backoff));
    }
  }
}

CLOSE THE AI. EXPLAIN THIS.

If five layers each make up to three attempts, how many calls can one user request cause at the bottom, and which layer should keep its retries?

HOW IT WORKS UNDERNEATH

SOURCES

Guide reviewed