Zum Inhalt springen
Beta

REL 6. How do you degrade gracefully and shed load deliberately?

Zuletzt aktualisiert am

A system that fails completely when one dependency is unavailable has treated every component as essential. Usually only some are, and the difference between an incident and an outage is whether anyone decided which.

The decisions here are made at design time. Under pressure, in the middle of an incident, nobody will work out which features are droppable, so a workload without a degradation plan does not degrade. It stops.

  • REL 6.1 Decide in advance which functions may be dropped, and in what order
  • REL 6.2 Serve reduced or stale results rather than errors where correctness allows
  • REL 6.3 Shed load at the edge before the system saturates
  • REL 6.4 Make the degraded state visible to operators and to users

REL 6.1 Decide in advance which functions may be dropped, and in what order

Section titled “REL 6.1 Decide in advance which functions may be dropped, and in what order”

Risk if not established: High

Take the flows from REL 2 and the failure modes from REL 3, and for each combination answer one question: can this flow continue without this component, and what does the user lose?

The answers cluster into a small set:

  • Essential. The flow cannot proceed. A checkout without payment processing is not a checkout.
  • Enrichment. The flow proceeds with less. Recommendations, personalization, related items.
  • Deferrable. The work must happen but not now. Confirmation emails, analytics events, index updates.
  • Cosmetic. Nobody notices. Usage counters, non-critical badges.

Most teams find that considerably more of their system is enrichment or deferrable than they assumed, which is the useful output of the exercise.

Order the drops. Under increasing pressure you want to shed cosmetic first, then enrichment, then defer what can be deferred, and only then fail. An ordering agreed in advance is what allows this to be automatic.

Deferring is not the same as dropping. If work is deferred it needs somewhere to go and something to drain it later, and that queue is now a component with its own failure modes.

On STACKIT. The decision is entirely yours; no platform feature identifies which of your features are essential.

Deferral usually needs a durable queue. RabbitMQ is the managed option, and Object Storage works well for deferred bulk work. Both introduce a component that must be as available as the deferral strategy assumes, which is worth composing into REL 1.3 rather than treating as free.

Tradeoffs. Operational Excellence. Each degradation path is a code path that exists for rare conditions, which means it is rarely executed and therefore rarely correct unless deliberately tested. REL 10.3 is what keeps them honest.

Verify. List the components on your most critical flow. Which are essential, and which have a defined behaviour when they are unavailable? When was that behaviour last executed?


REL 6.2 Serve reduced or stale results rather than errors where correctness allows

Section titled “REL 6.2 Serve reduced or stale results rather than errors where correctness allows”

Risk if not established: Medium

A cached value from ten minutes ago is usually better than an error page. Not always, and the distinction is a correctness question rather than a preference.

Stale is acceptable when the consequence of acting on old data is small: a product description, a configuration value, a list of categories. It is unacceptable where the data governs a decision: account balances, permissions, inventory at the point of sale, anything a regulator would ask about.

Decide per data set, not per system, and write the decision down alongside the maximum tolerable staleness. That number is what lets a cache serve during a failure without anyone having to judge in the moment.

The dangerous version of this is accidental. A cache that serves stale data during an outage because nobody configured what happens when the origin fails has produced graceful degradation by luck, and it will produce a correctness bug the same way. The mechanism is identical; only the intent differs, and only one of them is safe.

On STACKIT. Caching is an application concern. Managed building blocks include the Key Value Store , Redis , and CDN distributions for content served at the edge.

What matters is not which cache you use but whether its behaviour on origin failure was configured deliberately. Serving stale on error is usually an explicit setting, and the default is frequently not what a degradation plan wants.

Tradeoffs. Security. A cache holds a copy of the data with the same classification, which SEC 3 and SOV 3 both apply to. Cost Optimization. Caching trades storage and memory for availability and latency, and the trade is worth re-examining as traffic changes, per PERF 9.

Verify. For your most-read data set, what happens when the origin is unavailable: error, stale value, or empty result? Was that chosen, and what is the maximum staleness anyone agreed to?


REL 6.3 Shed load at the edge before the system saturates

Section titled “REL 6.3 Shed load at the edge before the system saturates”

Risk if not established: High

A system past its capacity does not fail gracefully on its own. Queues grow, latency rises, timeouts fire, retries add load, and throughput collapses well below what the system could have sustained. Total failure under overload is the normal outcome, not the exceptional one.

Deliberate shedding avoids this by rejecting some work early so the rest completes. Rejecting ten percent of requests quickly at the edge is a far better outcome than accepting all of them and serving none.

Two design points decide whether it works:

Reject early and cheaply. Shedding after a request has consumed database connections has already spent what you were protecting. The edge is where it belongs.

Shed by value, not at random. The flow ranking from REL 2.2 is the input. Background jobs and low-value traffic go first; the checkout path goes last. Uniform shedding treats a bulk export the same as a paying customer.

Rate limits are the common form and deserve a caveat: a limit tuned for abuse prevention is not the same as a limit tuned for capacity protection, and using one for both usually serves neither well.

On STACKIT. Application Load Balancer with WAF is the edge component where rules of this kind are usually applied, and in Kubernetes Engine an ingress controller or gateway you operate gives you the same position.

Aus der STACKIT-DokuALB WAF Funktionen › Funktionsübersicht und VerfügbarkeitStand der Quelle 19.08.2026 · übernommen 05.10.2026

Die folgende Tabelle listet jede Funktions-Gruppe auf und zeigt, wo sie heute verfügbar ist. Zeilen mit WIP zeigen an, dass die Schnittstelle noch entwickelt wird und noch nicht stabil für den Produktions-Einsatz ist.

Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

How much shedding happens at the edge rather than in application code depends on what your load balancer and WAF expose, and on whether they can tell traffic apart by path or by client. Establish that for the edge you actually run before designing against it, and keep the application-side path viable until you have.

Resource limits inside Kubernetes are a related but different control: they bound what a workload consumes, which protects neighbours rather than shedding load in a way the caller can respond to.

Tradeoffs. Performance Efficiency. A shedding threshold set too low rejects work the system could have handled. Tune against observed saturation points, which requires the baseline from PERF 2.

Verify. At what load does your workload start rejecting requests, and is that threshold chosen or emergent? Which traffic is shed first?


REL 6.4 Make the degraded state visible to operators and to users

Section titled “REL 6.4 Make the degraded state visible to operators and to users”

Risk if not established: Medium

Silent degradation is a trap. The system continues, dashboards look acceptable, and nobody investigates while a feature has been quietly broken for a week. Worse, the degradation can mask the failure it is compensating for, so the underlying problem is never fixed.

Operators need a signal that is distinct from healthy and from failed. A flow serving stale data because its origin is down is neither, and a health model that only has two states cannot express it. This connects directly to REL 10.1.

Users need to know when what they are seeing is incomplete, in proportion to the consequence. A missing recommendations panel needs no explanation. A stale balance does.

Add a duration bound. Degradation is a bridge to recovery, not a resting state. If a flow has been degraded for longer than the RTO from REL 1.2, that is an incident regardless of whether anything is technically failing.

On STACKIT. Observability is where degraded-state metrics and the alerts on them live, using alert groups to separate “degraded” from “down” so the two get different responses. Emitting the signal in the first place is application instrumentation, which is OPS 7.

For the platform’s own side of this, status.stackit.cloud reports service incidents, which is often the first place to check when your degradation triggered without an obvious cause of your own.

Tradeoffs. Little. The main cost is designing a health model with more than two states, which REL 10.1 requires anyway.

Verify. When your workload is running degraded, which alert fires and how does it differ from the alert for a full outage? How long can it stay degraded before someone is required to act?


  • REL 2 Critical flows, which supplies the ranking for what to shed
  • REL 3 Failure modes, where degradation is chosen as the response
  • REL 5 Resilient interactions, particularly what happens when a circuit opens
  • REL 10 Health model and testing, which must express degraded as a state
  • PERF 2 Baseline, which tells you where saturation actually begins