Requirement: the client waits at most 200 ms for a response.
A fixed timeout on every downstream call does not meet the requirement, because a fixed timeout bounds only one call. The fix is to compute each call timeout at call time, from the time left until one deadline that the whole request shares.
Example
We have three dependencies:
| Dependency | p50 | p99.9 | Required |
|---|---|---|---|
A, user profile | 10 ms | 40 ms | Yes |
B, pricing | 30 ms | 90 ms | Yes |
C, recommendations | 25 ms | 120 ms | No, has a fallback |
Our service calls A, then B, then C, because each call needs the result of the previous one. Our own work takes 20 ms.
Why fixed timeouts fail
Split the 200 ms in advance, so that the fixed timeouts add up to exactly 200 ms:
Agets 40 ms.Bgets 100 ms.Cgets 40 ms.- Own work gets 20 ms.
The split fails in two ways:
- A call cannot use time that an earlier call did not use.
Aanswers in 10 ms, andBneeds 105 ms, 5 ms more than its timeout.Btimes out, and the request fails. At that moment 90 ms of the 200 ms are still left, andCand our own work need at most 60 ms of them.
- A retry does not fit. A retry of
Bneeds another 100 ms that the split does not have.
Set one deadline per request
Both failures have one cause: each timeout is fixed before the request starts. The fix is to compute each call timeout when the call starts, from the time left. To know the time left, our service must know when the client stops waiting. That point in time is the deadline.
- The first service that receives the client’s request sets the deadline once. The deadline is the arrival time plus 200 ms.
- The caller sets the call timeout on each call. The called service receives it as its own deadline and computes its own call timeouts from it.
Compute each call timeout from the time left
The caller computes the timeout of each call from the time left when it makes the call:
The reserve is the time our service needs after the call returns: merging, serializing, and the trip back to the caller.
Per-dependency cap
The per-dependency cap is the longest time we wait for one dependency, even when more time is left. Pick it from an acceptable rate of false timeouts, timeouts on calls that would have succeeded. Amazon picks a rate such as 0.1% and uses the matching latency percentile, p99.9.
In this example, the caps are:
Agets 50 ms, its p99.9 of 40 ms plus some extra time.Bgets 110 ms, its p99.9 of 90 ms plus some extra time.Cgets 60 ms, less than its p99.9 of 120 ms, becauseChas a fallback. We accept more false timeouts onCso that the page loads sooner.
Minimum timeout
If the call timeout is too small for the dependency to reply, the call will probably time out. Give each dependency a minimum timeout, for example its p50 latency. If the call timeout is less than the minimum timeout, do not make the call. Use the fallback, or return an error if the dependency is required.
Apply the formula to a chain of calls
Each call needs the result of the previous one, so the reserve also includes minimum timeouts (p50 in our example) for all the required calls still to come. Do not reserve time for optional calls.
Take the chain from “Why fixed timeouts fail” with the same latencies.
Astarts at 0 ms. The reserve is 20 ms (own work) plusB’s minimum timeout of 30 ms, total 50 ms.Agets min(200 − 0 − 50, 50) = 50 ms and answers in 10 ms.Bstarts at 10 ms.Cis optional, so the reserve is only 20 ms.Bgets min(200 − 10 − 20, 110) = 110 ms and answers in 105 ms. With the fixed split,Bwould time out here.Cstarts at 115 ms and gets min(200 − 115 − 20, 60) = 60 ms. It answers in 25 ms, and our service responds at 160 ms.
The caps of A, B and C add up to 220 ms, and our own work needs 20 ms more. A fixed split cannot give these timeouts, because its shares must add up to 200 ms or less. The formula can use these caps, because a call gets its full cap only if enough time is left. When the earlier calls are slow, the later calls get shorter timeouts.
For example, if A answers in 48 ms and B in 108 ms, C starts at 156 ms and gets min(200 − 156 − 20, 60) = 24 ms. That is less than C’s minimum timeout of 25 ms, so our service skips C and responds at 176 ms.
Retry and hedge within the deadline
Retries. Before each retry, compute the call timeout again and compare it with the minimum timeout. No retry runs after the deadline.
Hedged requests. For a slow call, a retry starts only when the call timeout has passed, and by then little time is left. A hedged request does not wait for the timeout. If the original call has waited longer than the dependency’s p95 latency, send a second copy with the call timeout computed from the time left. Use whichever answer arrives first, and cancel the other call. This adds about 5% more requests (see The Tail at Scale).