Zum Inhalt springen
Beta

Reliability

Zuletzt aktualisiert am

Does it keep working when something breaks?

Reliability is the workload’s ability to deliver its promised function in the presence of failure. Not the absence of failure, which is not on offer. Disks fail, zones lose power, dependencies return errors, and someone deploys a bad configuration on a Friday afternoon. Reliability is what you have designed to happen next.

  • Deciding what “reliable enough” means for this workload, in business terms
  • Identifying which flows matter and ranking them
  • Understanding how the workload can fail, and deciding what happens when it does
  • Redundancy, resilience patterns, and graceful degradation
  • Backup, restore, and disaster recovery
  • Proving all of the above through testing rather than assuming it

Reliability against deliberate disruption belongs to Security, a denial-of- service attack and a traffic spike look similar and are addressed differently. The tooling and process for detecting and responding to incidents belongs to Operational Excellence; this pillar defines what needs detecting. Whether the system is fast enough belongs to Performance Efficiency, though the two meet at the point where saturation becomes an outage.

Reliability is bought, not designed in for free. Every nine costs money, complexity, and operational load, and the cost is not linear. Which is why this pillar begins with a business conversation rather than a technical one: until you know what an hour of downtime costs and how much data loss the business can absorb, any architectural decision about redundancy is a guess dressed as engineering.

The second idea, which practitioners reach later than they expect: recovery matters more than prevention. Effort spent making failure less likely has a ceiling; effort spent making recovery fast, routine, and rehearsed does not. A workload that fails twice a year and recovers in ninety seconds is more reliable in every way the business cares about than one that fails once a year and takes six hours to come back.

  1. Design principles: the reasoning, in five short statements
  2. Tradeoffs: what pursuing reliability costs the other pillars

Ten questions. Numbers follow the order the decisions are usually made in and do not indicate priority.

Each links to its full page: the reasoning, what STACKIT provides, the tradeoffs, and the assessment questions.


Establish availability targets, RTO, and RPO per flow, based on what an outage and what data loss actually cost the business. Record who agreed to them.

→ Best practices

Enumerate the paths through the workload that deliver business value, and rank them by consequence of failure. Reliability investment follows the ranking; not all components deserve equal treatment.

→ Best practices

REL 3: How do you analyse failure modes and decide what happens when each occurs?

Section titled “REL 3: How do you analyse failure modes and decide what happens when each occurs?”

For each component and dependency on a critical flow, identify how it can fail, what the effect is, and what the system does about it. An identified failure mode with no assigned mitigation is an accepted risk and must be recorded as one.

→ Best practices

Distribute components across availability zones, and across regions where the targets require it. Redundancy must cover state, not only compute, and the failover path itself must not become the single point of failure.

→ Best practices

Every call that crosses a process boundary needs a timeout and a retry policy with backoff and jitter; add a circuit breaker where a durably unhealthy dependency would otherwise consume the caller. Retried operations must be idempotent, or retries turn a transient fault into data corruption.

→ Best practices

Decide in advance which functions may be dropped when a dependency fails or demand exceeds capacity. Partial service is almost always better than total failure, and load shed on purpose is better than a system that collapses under its own queue.

→ Best practices

REL 7: How do you scale to absorb demand, and how much headroom do you keep?

Section titled “REL 7: How do you scale to absorb demand, and how much headroom do you keep?”

Size for realistic peaks, automate scaling where the workload allows it, and keep enough headroom to survive both the growth you predicted and the spike you did not. Know where the scaling limit is before you meet it.

→ Best practices

REL 8: How do you back up data, and how do you know you can restore it?

Section titled “REL 8: How do you back up data, and how do you know you can restore it?”

Back up every stateful component to a schedule derived from its RPO, hold copies where a single failure cannot destroy both, and restore from them on a regular cadence into a clean environment. An untested backup is not a backup.

→ Best practices

Document what happens when a whole region, a whole platform service, or the primary data set is lost: who decides, what the sequence is, what the dependencies are. Rehearse it, and time the rehearsal against the RTO you committed to.

→ Best practices

REL 10: How do you know the workload is healthy, and how do you test that it stays so?

Section titled “REL 10: How do you know the workload is healthy, and how do you test that it stays so?”

Define what “healthy” means per flow, in terms that map signals to a verdict rather than to a dashboard. Then verify continuously through fault injection, dependency failure simulation, and failover drills, in production where you can, in a production-like environment where you cannot.

→ Best practices