---
id: REL06
pillar: reliability
title: REL 6. How do you degrade gracefully and shed load deliberately?
description: Partial service beats total failure, and load shed on purpose beats a system collapsing under its own queue. How to decide what may be dropped, in advance.
status: draft
services: [kubernetes-engine, object-storage]
sidebar:
  order: 15
  label: Graceful degradation
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/rel-06-graceful-degradation/"
source_file: "docs/architecture/pillars/reliability/rel-06-graceful-degradation.mdx"
---

A system that fails completely when one dependency is unavailable has treated every component as
essential. Usually only some are, and the difference between an incident and an outage is whether
anyone decided which.

The decisions here are made at design time. Under pressure, in the middle of an incident, nobody
will work out which features are droppable, so a workload without a degradation plan does not
degrade. It stops.

## Best practices

- [`REL 6.1`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-61-decide-in-advance-which-functions-may-be-dropped-and-in-what-order) Decide in advance which functions may be dropped, and in what order
- [`REL 6.2`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-62-serve-reduced-or-stale-results-rather-than-errors-where-correctness-allows) Serve reduced or stale results rather than errors where correctness allows
- [`REL 6.3`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-63-shed-load-at-the-edge-before-the-system-saturates) Shed load at the edge before the system saturates
- [`REL 6.4`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-64-make-the-degraded-state-visible-to-operators-and-to-users) Make the degraded state visible to operators and to users

---

## REL 6.1 Decide in advance which functions may be dropped, and in what order

**Risk if not established:** High

Take the flows from [`REL 2`](/architecture/pillars/reliability/rel-02-critical-flows/) and the failure modes from [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/), and for each combination answer
one question: can this flow continue without this component, and what does the user lose?

The answers cluster into a small set:

- **Essential.** The flow cannot proceed. A checkout without payment processing is not a checkout.
- **Enrichment.** The flow proceeds with less. Recommendations, personalization, related items.
- **Deferrable.** The work must happen but not now. Confirmation emails, analytics events, index
  updates.
- **Cosmetic.** Nobody notices. Usage counters, non-critical badges.

Most teams find that considerably more of their system is enrichment or deferrable than they
assumed, which is the useful output of the exercise.

Order the drops. Under increasing pressure you want to shed cosmetic first, then enrichment, then
defer what can be deferred, and only then fail. An ordering agreed in advance is what allows this
to be automatic.

Deferring is not the same as dropping. If work is deferred it needs somewhere to go and something
to drain it later, and that queue is now a component with its own failure modes.

**On STACKIT.** The decision is entirely yours; no platform feature identifies which of your
features are essential.

Deferral usually needs a durable queue.
<LinkChip href="https://docs.stackit.cloud/products/messaging/rabbitmq/">RabbitMQ</LinkChip> is the managed option, and
<LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/">Object Storage</LinkChip> works well for
deferred bulk work. Both introduce a component that must be as available as the deferral strategy
assumes, which is worth composing into [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on) rather than treating as free.

**Tradeoffs.** **Operational Excellence.** Each degradation path is a code path that exists for
rare conditions, which means it is rarely executed and therefore rarely correct unless
deliberately tested. [`REL 10.3`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-103-inject-faults-and-verify-the-response-matches-the-design) is what keeps them honest.

**Verify.** List the components on your most critical flow. Which are essential, and which have a
defined behaviour when they are unavailable? When was that behaviour last executed?

---

## REL 6.2 Serve reduced or stale results rather than errors where correctness allows

**Risk if not established:** Medium

A cached value from ten minutes ago is usually better than an error page. Not always, and the
distinction is a correctness question rather than a preference.

Stale is acceptable when the consequence of acting on old data is small: a product description, a
configuration value, a list of categories. It is unacceptable where the data governs a decision:
account balances, permissions, inventory at the point of sale, anything a regulator would ask
about.

Decide per data set, not per system, and write the decision down alongside the maximum tolerable
staleness. That number is what lets a cache serve during a failure without anyone having to judge
in the moment.

The dangerous version of this is accidental. A cache that serves stale data during an outage
because nobody configured what happens when the origin fails has produced graceful degradation by
luck, and it will produce a correctness bug the same way. The mechanism is identical; only the
intent differs, and only one of them is safe.

**On STACKIT.** Caching is an application concern. Managed building blocks include the
<LinkChip href="https://docs.stackit.cloud/products/databases/key-value-store/">Key Value Store</LinkChip>,
<LinkChip href="https://docs.stackit.cloud/products/databases/redis/">Redis</LinkChip>, and
<LinkChip href="https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/">CDN distributions</LinkChip>
for content served at the edge.

What matters is not which cache you use but whether its behaviour on origin failure was configured
deliberately. Serving stale on error is usually an explicit setting, and the default is frequently
not what a degradation plan wants.

**Tradeoffs.** **Security.** A cache holds a copy of the data with the same classification, which
[`SEC 3`](/architecture/pillars/security/sec-03-data-classification/) and [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/) both apply to. **Cost Optimization.** Caching trades storage and memory for
availability and latency, and the trade is worth re-examining as traffic changes, per [`PERF 9`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/).

**Verify.** For your most-read data set, what happens when the origin is unavailable: error, stale
value, or empty result? Was that chosen, and what is the maximum staleness anyone agreed to?

---

## REL 6.3 Shed load at the edge before the system saturates

**Risk if not established:** High

A system past its capacity does not fail gracefully on its own. Queues grow, latency rises,
timeouts fire, retries add load, and throughput collapses well below what the system could have
sustained. Total failure under overload is the normal outcome, not the exceptional one.

Deliberate shedding avoids this by rejecting some work early so the rest completes. Rejecting ten
percent of requests quickly at the edge is a far better outcome than accepting all of them and
serving none.

Two design points decide whether it works:

**Reject early and cheaply.** Shedding after a request has consumed database connections has
already spent what you were protecting. The edge is where it belongs.

**Shed by value, not at random.** The flow ranking from [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume) is the input. Background jobs
and low-value traffic go first; the checkout path goes last. Uniform shedding treats a bulk export
the same as a paying customer.

Rate limits are the common form and deserve a caveat: a limit tuned for abuse prevention is not
the same as a limit tuned for capacity protection, and using one for both usually serves neither
well.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/">Application Load Balancer with
WAF</LinkChip>
is the edge component where rules of this kind are usually applied, and in <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/">Kubernetes
Engine</LinkChip> an ingress controller or
gateway you operate gives you the same position.

> From the STACKIT docs: [ALB WAF features › Feature overview and availability](https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/#feature-overview-and-availability) (Source updated 06.10.2026, copied 06.10.2026)

The following table lists each feature area and where it is available today. Rows marked `WIP` indicate that the interface is being built and is not yet stable for production use.

| Feature area | API | SDK | Terraform | Portal |
| --- | --- | --- | --- | --- |
| [ALB WAF configuration](https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/#alb-waf-configuration) | Available | Available | Available | Available |
| [Listener attachment](https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/#listener-attachment) | Available | Available | Available | Available |
| [Managed rule sets (OWASP CRS)](https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/#managed-rule-sets) | Available | WIP | Available | Available |
| [Per-rule mode override (enable/disable/log-only)](https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/#per-rule-overrides) | Available | WIP | WIP | Available |
| [Custom rule groups (SecLang abstraction)](https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/#custom-rule-groups) | Available | WIP | Available | Available |
| [Project quotas for ALB WAF, MRS, and CRG](https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/basics/features-alb-waf/#quotas) | Available | WIP | WIP | WIP |

How much shedding happens at the edge rather than in application code depends on what your load
balancer and WAF expose, and on whether they can tell traffic apart by path or by client.
Establish that for the edge you actually run before designing against it, and keep the
application-side path viable until you have.

Resource limits inside Kubernetes are a related but different control: they bound what a workload
consumes, which protects neighbours rather than shedding load in a way the caller can respond to.

**Tradeoffs.** **Performance Efficiency.** A shedding threshold set too low rejects work the
system could have handled. Tune against observed saturation points, which requires the baseline
from [`PERF 2`](/architecture/pillars/performance-efficiency/perf-02-baseline/).

**Verify.** At what load does your workload start rejecting requests, and is that threshold chosen
or emergent? Which traffic is shed first?

---

## REL 6.4 Make the degraded state visible to operators and to users

**Risk if not established:** Medium

Silent degradation is a trap. The system continues, dashboards look acceptable, and nobody
investigates while a feature has been quietly broken for a week. Worse, the degradation can mask
the failure it is compensating for, so the underlying problem is never fixed.

Operators need a signal that is distinct from healthy and from failed. A flow serving stale data
because its origin is down is neither, and a health model that only has two states cannot express
it. This connects directly to [`REL 10.1`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-101-define-what-healthy-means-per-flow-mapping-signals-to-a-verdict).

Users need to know when what they are seeing is incomplete, in proportion to the consequence. A
missing recommendations panel needs no explanation. A stale balance does.

Add a duration bound. Degradation is a bridge to recovery, not a resting state. If a flow has been
degraded for longer than the RTO from [`REL 1.2`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-12-set-availability-rto-and-rpo-per-critical-flow-rather-than-per-workload), that is an incident regardless of whether
anything is technically failing.

**On STACKIT.**
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> is
where degraded-state metrics and the alerts on them live, using <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/manage-alert-groups-and-alerts/">alert
groups</LinkChip>
to separate "degraded" from "down" so the two get different responses. Emitting the signal in the
first place is application instrumentation, which is [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/).

For the platform's own side of this,
<LinkChip href="https://status.stackit.cloud">status.stackit.cloud</LinkChip> reports service incidents, which is often the
first place to check when your degradation triggered without an obvious cause of your own.

**Tradeoffs.** Little. The main cost is designing a health model with more than two states, which
[`REL 10.1`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-101-define-what-healthy-means-per-flow-mapping-signals-to-a-verdict) requires anyway.

**Verify.** When your workload is running degraded, which alert fires and how does it differ from
the alert for a full outage? How long can it stay degraded before someone is required to act?

---

## Related

- [`REL 2`](/architecture/pillars/reliability/rel-02-critical-flows/) Critical flows, which supplies the ranking for what to shed
- [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) Failure modes, where degradation is chosen as the response
- [`REL 5`](/architecture/pillars/reliability/rel-05-resilient-interactions/) Resilient interactions, particularly what happens when a circuit opens
- [`REL 10`](/architecture/pillars/reliability/rel-10-health-model-and-testing/) Health model and testing, which must express degraded as a state
- [`PERF 2`](/architecture/pillars/performance-efficiency/perf-02-baseline/) Baseline, which tells you where saturation actually begins
