---
id: REL05
pillar: reliability
title: REL 5. How do you make remote interactions resilient?
description: Timeouts, retries with backoff, idempotency and circuit breakers. The patterns that stop one slow dependency from turning into a workload-wide outage.
status: draft
services: [kubernetes-engine]
sidebar:
  order: 14
  label: Resilient interactions
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/rel-05-resilient-interactions/"
source_file: "docs/architecture/pillars/reliability/rel-05-resilient-interactions.mdx"
---

Every call that leaves a process is a call that can hang, fail, or return something unexpected.
Most outages that look like a single component failing are actually a single component becoming
slow, and everything that calls it waiting.

This question is almost entirely about your own code. The platform gives you places to enforce
some of it, but a missing timeout is a missing timeout regardless of where the workload runs.

## Best practices

- [`REL 5.1`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-51-set-an-explicit-timeout-on-every-call-that-crosses-a-process-boundary) Set an explicit timeout on every call that crosses a process boundary
- [`REL 5.2`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-52-retry-with-backoff-and-jitter-and-only-what-is-safe-to-retry) Retry with backoff and jitter, and only what is safe to retry
- [`REL 5.3`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-53-make-retried-operations-idempotent) Make retried operations idempotent
- [`REL 5.4`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-54-add-a-circuit-breaker-so-a-slow-dependency-does-not-become-an-outage) Add a circuit breaker so a slow dependency does not become an outage

---

## REL 5.1 Set an explicit timeout on every call that crosses a process boundary

**Risk if not established:** High

The default timeout of most clients is either very long or absent. A caller with no timeout will
wait indefinitely, holding a thread, a connection and whatever the caller above it was holding.
This is how a single slow dependency consumes an entire application tier.

Derive the timeout from the flow's latency target in [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/), not from how long the dependency
usually takes. A budget approach works well: the flow has a total time, each hop gets a share, and
a hop that exceeds its share has already failed the flow even if it eventually answers.

Set them explicitly everywhere, including the calls that feel safe. Database drivers, HTTP
clients, DNS resolution, connection establishment and TLS handshakes all have separate settings,
and the one you forget is the one that hangs.

Be aware of what a timeout does not do. Abandoning the wait does not abandon the work: the callee
may still complete the operation. That is the reason [`REL 5.3`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-53-make-retried-operations-idempotent) exists.

**On STACKIT.** This is application configuration rather than platform configuration, and no
provider can set it for you.

Where the platform helps is at the edges. Health checks and connection handling in
<LinkChip href="https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/">load balancing</LinkChip>
determine how long a request to an unhealthy backend hangs before it is taken out of rotation, and
in
<LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/">Kubernetes Engine</LinkChip> the readiness
and liveness probes you configure decide how quickly a struggling pod stops receiving traffic.
Neither replaces client-side timeouts; both shorten the window in which they matter.

**Tradeoffs.** **Performance Efficiency.** A timeout that is too aggressive fails requests that
would have succeeded, which converts a latency problem into an error-rate problem. Tune against
observed latency distributions rather than against averages.

**Verify.** Pick a service on your critical flow. List every outbound call it makes and the
timeout configured on each. Which of those are inherited defaults nobody chose?

---

## REL 5.2 Retry with backoff and jitter, and only what is safe to retry

**Risk if not established:** High

Retries turn transient faults into successes, and they turn overloaded dependencies into dead
ones. The difference is entirely in how they are configured.

Three properties are needed together:

**Backoff.** Exponentially increasing delay, so a struggling dependency gets less traffic rather
than more.

**Jitter.** Randomized delay, so retries from many callers do not arrive simultaneously. Without
it, backoff synchronizes clients into waves, which is worse than no backoff at all.

**A bound.** A maximum attempt count or a total time budget. Unbounded retries do not improve the
outcome; they extend the outage and hide it from the caller above.

Retry only what is worth retrying. A timeout or a 503 is a reasonable candidate. A validation
error or an authorization failure will fail identically every time, and retrying it only adds
load.

Watch for retry amplification. If three layers each retry three times, one user request becomes
twenty-seven calls to the bottom service, precisely when it is least able to take them. Retry at
one layer, usually the one closest to the failure.

**On STACKIT.** Retry behaviour lives in your application and your client libraries. The
<LinkChip href="https://docs.stackit.cloud/">STACKIT SDKs and CLI</LinkChip> have their own retry behaviour for API calls,
which applies when your automation talks to the platform but does not govern calls between your
own services.

Service quotas are worth noting here. Retrying against a quota limit will not succeed and consumes
the rate you have; treat quota errors as non-retryable and handle them as capacity signals under
[`REL 7.3`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-73-find-every-scaling-limit-before-you-approach-it) instead.

**Tradeoffs.** **Performance Efficiency.** Retries add load at the worst moment, which is why the
bound and the circuit breaker in [`REL 5.4`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-54-add-a-circuit-breaker-so-a-slow-dependency-does-not-become-an-outage) matter. **Cost Optimization.** Amplified retries
against a metered service are billable.

**Verify.** For one critical dependency, what is the retry policy: how many attempts, what
backoff, is there jitter, and which error classes are excluded? How many layers of your stack
retry the same call?

---

## REL 5.3 Make retried operations idempotent

**Risk if not established:** High

A timeout means you do not know whether the operation happened. Retrying a non-idempotent
operation in that state is how a transient network fault becomes a duplicate payment.

This is a correctness requirement, not a reliability nicety, and it constrains the interface
rather than the client. Reads are naturally idempotent. Writes need to be made so, usually with a
client-generated key that the server uses to recognize and collapse repeats.

Design it into the API rather than adding it later. Retrofitting idempotency onto an interface
that already has callers means changing every caller, and the callers you do not control will not
change.

Two adjacent cases are frequently missed. Message consumers need the same property, because
at-least-once delivery is the norm and redelivery after a partial failure is routine. And
operations that are idempotent in isolation may not be in sequence: applying the same update twice
is safe, applying it after a later update is not.

**On STACKIT.** This is an application design property throughout. No platform feature provides
it, and no provider can.

Where the platform intersects is in managed messaging.
<LinkChip href="https://docs.stackit.cloud/products/messaging/rabbitmq/">RabbitMQ</LinkChip> delivers at least once under
the usual failure conditions, so consumers of a queue need to be idempotent for the same reason
retried HTTP callers do.

**Tradeoffs.** **Performance Efficiency.** Deduplication requires storing and checking keys, which
adds a lookup on the write path and state that must itself be managed. Usually small, occasionally
material on high-volume endpoints.

**Verify.** For the most consequential write operation on your critical flow, what happens if the
client sends it twice with the same payload? Is that behaviour tested?

---

## REL 5.4 Add a circuit breaker so a slow dependency does not become an outage

**Risk if not established:** Medium

Timeouts and retries handle the individual call. A circuit breaker handles the situation where a
dependency is durably unhealthy and every call is going to fail, so continuing to try wastes your
capacity and prolongs theirs.

The mechanism is simple: track failures against a dependency, stop calling it once they exceed a
threshold, fail fast for a period, then let a probe request through to test recovery. The value is
that the caller stays healthy while the callee is not, which is what keeps a partial failure
partial.

The failing behaviour matters more than the breaker. When the circuit is open, what does the flow
do? Return an error, serve a cached value, skip the enrichment step, queue the work for later.
That decision belongs to [`REL 6`](/architecture/pillars/reliability/rel-06-graceful-degradation/) and should be made before the breaker is added, or you have
simply moved where the failure appears.

Tune with care. A threshold too sensitive opens on normal variance and creates outages that were
not happening; too tolerant and the breaker never opens. Both failure modes are common and both
need production data to correct.

**On STACKIT.** Circuit breaking is an application or service-mesh concern, and there is no
managed service mesh, so on
<LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/">Kubernetes Engine</LinkChip> this is either
a library in your application or a mesh you install and operate yourself. Both are legitimate; the
second is a real operational commitment under [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/) and should be a deliberate choice rather
than a default.

**Tradeoffs.** **Operational Excellence.** Another mechanism with thresholds to tune and a state
machine to reason about during incidents. For workloads with few dependencies, timeouts and
bounded retries often suffice, and adding a breaker is complexity without benefit.

**Verify.** For one critical dependency, at what failure rate does the caller stop calling it, and
what does the flow do while the circuit is open? Has that path been exercised?

---

## Related

- [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) Failure modes, particularly the "slow" mode that this question addresses
- [`REL 6`](/architecture/pillars/reliability/rel-06-graceful-degradation/) Graceful degradation, which decides what happens when a call is abandoned
- [`REL 10`](/architecture/pillars/reliability/rel-10-health-model-and-testing/) Health model and testing, which is where these policies get exercised
- [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) Performance targets, which the timeout budget derives from
- [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) Observability, without which none of these policies can be tuned
