Zum Inhalt springen
Beta

REL 5. How do you make remote interactions resilient?

Zuletzt aktualisiert am

Every call that leaves a process is a call that can hang, fail, or return something unexpected. Most outages that look like a single component failing are actually a single component becoming slow, and everything that calls it waiting.

This question is almost entirely about your own code. The platform gives you places to enforce some of it, but a missing timeout is a missing timeout regardless of where the workload runs.

  • REL 5.1 Set an explicit timeout on every call that crosses a process boundary
  • REL 5.2 Retry with backoff and jitter, and only what is safe to retry
  • REL 5.3 Make retried operations idempotent
  • REL 5.4 Add a circuit breaker so a slow dependency does not become an outage

REL 5.1 Set an explicit timeout on every call that crosses a process boundary

Section titled “REL 5.1 Set an explicit timeout on every call that crosses a process boundary”

Risk if not established: High

The default timeout of most clients is either very long or absent. A caller with no timeout will wait indefinitely, holding a thread, a connection and whatever the caller above it was holding. This is how a single slow dependency consumes an entire application tier.

Derive the timeout from the flow’s latency target in PERF 1, not from how long the dependency usually takes. A budget approach works well: the flow has a total time, each hop gets a share, and a hop that exceeds its share has already failed the flow even if it eventually answers.

Set them explicitly everywhere, including the calls that feel safe. Database drivers, HTTP clients, DNS resolution, connection establishment and TLS handshakes all have separate settings, and the one you forget is the one that hangs.

Be aware of what a timeout does not do. Abandoning the wait does not abandon the work: the callee may still complete the operation. That is the reason REL 5.3 exists.

On STACKIT. This is application configuration rather than platform configuration, and no provider can set it for you.

Where the platform helps is at the edges. Health checks and connection handling in load balancing determine how long a request to an unhealthy backend hangs before it is taken out of rotation, and in Kubernetes Engine the readiness and liveness probes you configure decide how quickly a struggling pod stops receiving traffic. Neither replaces client-side timeouts; both shorten the window in which they matter.

Tradeoffs. Performance Efficiency. A timeout that is too aggressive fails requests that would have succeeded, which converts a latency problem into an error-rate problem. Tune against observed latency distributions rather than against averages.

Verify. Pick a service on your critical flow. List every outbound call it makes and the timeout configured on each. Which of those are inherited defaults nobody chose?


REL 5.2 Retry with backoff and jitter, and only what is safe to retry

Section titled “REL 5.2 Retry with backoff and jitter, and only what is safe to retry”

Risk if not established: High

Retries turn transient faults into successes, and they turn overloaded dependencies into dead ones. The difference is entirely in how they are configured.

Three properties are needed together:

Backoff. Exponentially increasing delay, so a struggling dependency gets less traffic rather than more.

Jitter. Randomized delay, so retries from many callers do not arrive simultaneously. Without it, backoff synchronizes clients into waves, which is worse than no backoff at all.

A bound. A maximum attempt count or a total time budget. Unbounded retries do not improve the outcome; they extend the outage and hide it from the caller above.

Retry only what is worth retrying. A timeout or a 503 is a reasonable candidate. A validation error or an authorization failure will fail identically every time, and retrying it only adds load.

Watch for retry amplification. If three layers each retry three times, one user request becomes twenty-seven calls to the bottom service, precisely when it is least able to take them. Retry at one layer, usually the one closest to the failure.

On STACKIT. Retry behaviour lives in your application and your client libraries. The STACKIT SDKs and CLI have their own retry behaviour for API calls, which applies when your automation talks to the platform but does not govern calls between your own services.

Service quotas are worth noting here. Retrying against a quota limit will not succeed and consumes the rate you have; treat quota errors as non-retryable and handle them as capacity signals under REL 7.3 instead.

Tradeoffs. Performance Efficiency. Retries add load at the worst moment, which is why the bound and the circuit breaker in REL 5.4 matter. Cost Optimization. Amplified retries against a metered service are billable.

Verify. For one critical dependency, what is the retry policy: how many attempts, what backoff, is there jitter, and which error classes are excluded? How many layers of your stack retry the same call?


REL 5.3 Make retried operations idempotent

Section titled “REL 5.3 Make retried operations idempotent”

Risk if not established: High

A timeout means you do not know whether the operation happened. Retrying a non-idempotent operation in that state is how a transient network fault becomes a duplicate payment.

This is a correctness requirement, not a reliability nicety, and it constrains the interface rather than the client. Reads are naturally idempotent. Writes need to be made so, usually with a client-generated key that the server uses to recognize and collapse repeats.

Design it into the API rather than adding it later. Retrofitting idempotency onto an interface that already has callers means changing every caller, and the callers you do not control will not change.

Two adjacent cases are frequently missed. Message consumers need the same property, because at-least-once delivery is the norm and redelivery after a partial failure is routine. And operations that are idempotent in isolation may not be in sequence: applying the same update twice is safe, applying it after a later update is not.

On STACKIT. This is an application design property throughout. No platform feature provides it, and no provider can.

Where the platform intersects is in managed messaging. RabbitMQ delivers at least once under the usual failure conditions, so consumers of a queue need to be idempotent for the same reason retried HTTP callers do.

Tradeoffs. Performance Efficiency. Deduplication requires storing and checking keys, which adds a lookup on the write path and state that must itself be managed. Usually small, occasionally material on high-volume endpoints.

Verify. For the most consequential write operation on your critical flow, what happens if the client sends it twice with the same payload? Is that behaviour tested?


REL 5.4 Add a circuit breaker so a slow dependency does not become an outage

Section titled “REL 5.4 Add a circuit breaker so a slow dependency does not become an outage”

Risk if not established: Medium

Timeouts and retries handle the individual call. A circuit breaker handles the situation where a dependency is durably unhealthy and every call is going to fail, so continuing to try wastes your capacity and prolongs theirs.

The mechanism is simple: track failures against a dependency, stop calling it once they exceed a threshold, fail fast for a period, then let a probe request through to test recovery. The value is that the caller stays healthy while the callee is not, which is what keeps a partial failure partial.

The failing behaviour matters more than the breaker. When the circuit is open, what does the flow do? Return an error, serve a cached value, skip the enrichment step, queue the work for later. That decision belongs to REL 6 and should be made before the breaker is added, or you have simply moved where the failure appears.

Tune with care. A threshold too sensitive opens on normal variance and creates outages that were not happening; too tolerant and the breaker never opens. Both failure modes are common and both need production data to correct.

On STACKIT. Circuit breaking is an application or service-mesh concern, and there is no managed service mesh, so on Kubernetes Engine this is either a library in your application or a mesh you install and operate yourself. Both are legitimate; the second is a real operational commitment under OPS 10 and should be a deliberate choice rather than a default.

Tradeoffs. Operational Excellence. Another mechanism with thresholds to tune and a state machine to reason about during incidents. For workloads with few dependencies, timeouts and bounded retries often suffice, and adding a breaker is complexity without benefit.

Verify. For one critical dependency, at what failure rate does the caller stop calling it, and what does the flow do while the circuit is open? Has that path been exercised?


  • REL 3 Failure modes, particularly the “slow” mode that this question addresses
  • REL 6 Graceful degradation, which decides what happens when a call is abandoned
  • REL 10 Health model and testing, which is where these policies get exercised
  • PERF 1 Performance targets, which the timeout budget derives from
  • OPS 7 Observability, without which none of these policies can be tuned