Skip to content
bacebu4
Go back

Designing Deadlines and Timeouts

Читать на русском

Requirement: the client waits at most 200 ms for a response.

A fixed timeout on every downstream call does not meet the requirement, because a fixed timeout bounds only one call. The fix is to compute each call timeout at call time, from the time left until one deadline that the whole request shares.

Example

We have three dependencies:

Dependencyp50p99.9Required
A, user profile10 ms40 msYes
B, pricing30 ms90 msYes
C, recommendations25 ms120 msNo, has a fallback

Our service calls A, then B, then C, because each call needs the result of the previous one. Our own work takes 20 ms.

Why fixed timeouts fail

Split the 200 ms in advance, so that the fixed timeouts add up to exactly 200 ms:

The split fails in two ways:

Fixed split: the request fails while 90 ms are still left

Set one deadline per request

Both failures have one cause: each timeout is fixed before the request starts. The fix is to compute each call timeout when the call starts, from the time left. To know the time left, our service must know when the client stops waiting. That point in time is the deadline.

Compute each call timeout from the time left

The caller computes the timeout of each call from the time left when it makes the call:

call timeout=min⁡(deadline−now−reserve, per-dependency cap)\text{call timeout} = \min(\text{deadline} - \text{now} - \text{reserve},\ \text{per-dependency cap})

The reserve is the time our service needs after the call returns: merging, serializing, and the trip back to the caller.

Per-dependency cap

The per-dependency cap is the longest time we wait for one dependency, even when more time is left. Pick it from an acceptable rate of false timeouts, timeouts on calls that would have succeeded. Amazon picks a rate such as 0.1% and uses the matching latency percentile, p99.9.

In this example, the caps are:

Minimum timeout

If the call timeout is too small for the dependency to reply, the call will probably time out. Give each dependency a minimum timeout, for example its p50 latency. If the call timeout is less than the minimum timeout, do not make the call. Use the fallback, or return an error if the dependency is required.

Apply the formula to a chain of calls

Each call needs the result of the previous one, so the reserve also includes minimum timeouts (p50 in our example) for all the required calls still to come. Do not reserve time for optional calls.

Take the chain from “Why fixed timeouts fail” with the same latencies.

Same latencies: a timeout from the deadline lets B succeed

The caps of A, B and C add up to 220 ms, and our own work needs 20 ms more. A fixed split cannot give these timeouts, because its shares must add up to 200 ms or less. The formula can use these caps, because a call gets its full cap only if enough time is left. When the earlier calls are slow, the later calls get shorter timeouts.

For example, if A answers in 48 ms and B in 108 ms, C starts at 156 ms and gets min(200 − 156 − 20, 60) = 24 ms. That is less than C’s minimum timeout of 25 ms, so our service skips C and responds at 176 ms.

Slow A and B: with C, the response would come after the deadline, so C is skipped

Retry and hedge within the deadline

Retries. Before each retry, compute the call timeout again and compare it with the minimum timeout. No retry runs after the deadline.

Hedged requests. For a slow call, a retry starts only when the call timeout has passed, and by then little time is left. A hedged request does not wait for the timeout. If the original call has waited longer than the dependency’s p95 latency, send a second copy with the call timeout computed from the time left. Use whichever answer arrives first, and cancel the other call. This adds about 5% more requests (see The Tail at Scale).

A slow call to B: retry vs hedged request


Share this post:

Previous Post
When Your Saga Rollback Succeeds and Your Users Lose Money