A brief blip turns into a long outage
“A dependency hiccups for seconds, but traffic to it multiplies, error rates stay high after the root cause is gone, and it only recovers when load is shed manually.”
ERROR CODES YOU MAY SEE
A code is a clue. Use its meaning and surrounding evidence to narrow the cause.
503HTTP status codes · RFC 9110 · standardThe server is temporarily unable to handle the request due to overload or maintenance. It may send Retry-After.
Treat it as transient: honor Retry-After when present, otherwise back off with jitter and cap total attempts.
Official reference for 503 (opens in new tab)Code definition reviewedDEADLINE_EXCEEDEDgRPC status codes · standardThe deadline expired before the operation could complete (status 4). For state-changing operations it may be returned even if the operation completed.
Check whether the deadline was spent waiting in queues or on retries downstream. Propagate the caller's remaining deadline instead of starting a fresh timeout at each hop, and only retry idempotent calls.
Official reference for DEADLINE_EXCEEDED (opens in new tab)Code definition reviewed
THE PRINCIPLE
A retry is extra load. Every retry policy needs a cap, a budget, jitter, and a deadline it cannot outlive.
FIRST MOVES
- Map who retries: browser, SDK, gateway, service mesh, each service, the DB driver. Three layers of 3 retries is up to 64 attempts per user action.
- Retry at one layer only (usually closest to the caller that owns the intent); set the others to zero.
- Use capped exponential backoff with jitter, and a retry budget/token bucket so retries stop when most calls fail.
- Propagate one absolute deadline downstream; drop work whose caller has already given up.
- Shed load early and cheaply (429/503 with Retry-After) and add circuit breakers so callers fail fast instead of queueing.
- Retry only idempotent operations or ones carrying an idempotency key.
TOOLS: THEN → NOW
PATTERN SNAPSHOT
// The only retry layer: capped attempts, full jitter, shared deadline.
async function call<T>(fn: (s: AbortSignal) => Promise<T>, deadline: number) {
for (let attempt = 0; ; attempt++) {
try {
return await fn(AbortSignal.timeout(Math.max(0, deadline - Date.now())));
} catch (err) {
if (attempt >= 2 || !isRetryable(err) || !retryBudget.tryAcquire()) throw err;
const backoff = Math.random() * Math.min(2_000, 100 * 2 ** attempt);
if (Date.now() + backoff >= deadline) throw err;
await new Promise((r) => setTimeout(r, backoff));
}
}
}CLOSE THE AI. EXPLAIN THIS.
If five layers each make up to three attempts, how many calls can one user request cause at the bottom, and which layer should keep its retries?HOW IT WORKS UNDERNEATH
RELATED SYMPTOMS
SOURCES
Guide reviewed