---
type: pillar
pillar: reliability
code: REL
title: Reliability
description: Does it keep working when something breaks? Failure modes, redundancy, recovery and resilient interactions, and the testing that proves any of it works.
status: draft
sidebar:
  label: "Overview"
  order: 0
hideLinkCard: true
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/"
source_file: "docs/architecture/pillars/reliability/index.mdx"
---

> **Does it keep working when something breaks?**

Reliability is the workload's ability to deliver its promised function in the presence of failure.
Not the absence of failure, which is not on offer. Disks fail, zones lose power, dependencies
return errors, and someone deploys a bad configuration on a Friday afternoon. Reliability is what
you have designed to happen next.

## What this pillar covers

- Deciding what "reliable enough" means for this workload, in business terms
- Identifying which flows matter and ranking them
- Understanding how the workload can fail, and deciding what happens when it does
- Redundancy, resilience patterns, and graceful degradation
- Backup, restore, and disaster recovery
- Proving all of the above through testing rather than assuming it

## What it does not cover

Reliability against *deliberate* disruption belongs to [Security](/architecture/pillars/security/), a denial-of-
service attack and a traffic spike look similar and are addressed differently. The tooling and
process for detecting and responding to incidents belongs to [Operational
Excellence](/architecture/pillars/operational-excellence/); this pillar defines what needs detecting. Whether the
system is fast enough belongs to [Performance Efficiency](/architecture/pillars/performance-efficiency/), though the
two meet at the point where saturation becomes an outage.

## The central idea

Reliability is bought, not designed in for free. Every nine costs money, complexity, and
operational load, and the cost is not linear. Which is why this pillar begins with a business
conversation rather than a technical one: until you know what an hour of downtime costs and how
much data loss the business can absorb, any architectural decision about redundancy is a guess
dressed as engineering.

The second idea, which practitioners reach later than they expect: **recovery matters more than
prevention**. Effort spent making failure less likely has a ceiling; effort spent making recovery
fast, routine, and rehearsed does not. A workload that fails twice a year and recovers in ninety
seconds is more reliable in every way the business cares about than one that fails once a year and
takes six hours to come back.

## Where to start

1. [Design principles](/architecture/pillars/reliability/principles/): the reasoning, in five short statements
2. [Tradeoffs](/architecture/pillars/reliability/tradeoffs/): what pursuing reliability costs the other pillars

## Questions

Ten questions. Numbers follow the order the decisions are usually made in and do **not** indicate
priority.

Each links to its full page: the reasoning, what STACKIT provides, the tradeoffs, and the
assessment questions.

---

### [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/): How do you derive reliability targets from business impact?

Establish availability targets, RTO, and RPO per flow, based on what an outage and what data loss
actually cost the business. Record who agreed to them.

→ [Best practices](/architecture/pillars/reliability/rel-01-reliability-targets/)

### [`REL 2`](/architecture/pillars/reliability/rel-02-critical-flows/): How do you identify and rank the critical flows?

Enumerate the paths through the workload that deliver business value, and rank them by consequence
of failure. Reliability investment follows the ranking; not all components deserve equal
treatment.

→ [Best practices](/architecture/pillars/reliability/rel-02-critical-flows/)

### [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/): How do you analyse failure modes and decide what happens when each occurs?

For each component and dependency on a critical flow, identify how it can fail, what the effect
is, and what the system does about it. An identified failure mode with no assigned mitigation is
an accepted risk and must be recorded as one.

→ [Best practices](/architecture/pillars/reliability/rel-03-failure-mode-analysis/)

### [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/): How do you eliminate single points of failure?

Distribute components across availability zones, and across regions where the targets require it.
Redundancy must cover state, not only compute, and the failover path itself must not become the
single point of failure.

→ [Best practices](/architecture/pillars/reliability/rel-04-redundancy/)

### [`REL 5`](/architecture/pillars/reliability/rel-05-resilient-interactions/): How do you make remote interactions resilient?

Every call that crosses a process boundary needs a timeout and a retry policy with backoff and
jitter; add a circuit breaker where a durably unhealthy dependency would otherwise consume the
caller. Retried operations must be idempotent, or retries turn a transient fault into data
corruption.

→ [Best practices](/architecture/pillars/reliability/rel-05-resilient-interactions/)

### [`REL 6`](/architecture/pillars/reliability/rel-06-graceful-degradation/): How do you degrade gracefully and shed load deliberately?

Decide in advance which functions may be dropped when a dependency fails or demand exceeds
capacity. Partial service is almost always better than total failure, and load shed on purpose is
better than a system that collapses under its own queue.

→ [Best practices](/architecture/pillars/reliability/rel-06-graceful-degradation/)

### [`REL 7`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/): How do you scale to absorb demand, and how much headroom do you keep?

Size for realistic peaks, automate scaling where the workload allows it, and keep enough headroom
to survive both the growth you predicted and the spike you did not. Know where the scaling limit
is before you meet it.

→ [Best practices](/architecture/pillars/reliability/rel-07-scaling-and-headroom/)

### [`REL 8`](/architecture/pillars/reliability/rel-08-backup-and-restore/): How do you back up data, and how do you know you can restore it?

Back up every stateful component to a schedule derived from its RPO, hold copies where a single
failure cannot destroy both, and restore from them on a regular cadence into a clean environment.
An untested backup is not a backup.

→ [Best practices](/architecture/pillars/reliability/rel-08-backup-and-restore/)

### [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/): How do you plan and rehearse disaster recovery?

Document what happens when a whole region, a whole platform service, or the primary data set is
lost: who decides, what the sequence is, what the dependencies are. Rehearse it, and time the
rehearsal against the RTO you committed to.

→ [Best practices](/architecture/pillars/reliability/rel-09-disaster-recovery/)

### [`REL 10`](/architecture/pillars/reliability/rel-10-health-model-and-testing/): How do you know the workload is healthy, and how do you test that it stays so?

Define what "healthy" means per flow, in terms that map signals to a verdict rather than to a
dashboard. Then verify continuously through fault injection, dependency failure simulation, and
failover drills, in production where you can, in a production-like environment where you cannot.

→ [Best practices](/architecture/pillars/reliability/rel-10-health-model-and-testing/)

## Related

- <LinkChip href="/architecture/pillars/reliability/principles/">Design principles</LinkChip>
- <LinkChip href="/architecture/pillars/reliability/tradeoffs/">Tradeoffs</LinkChip>
