Reliability
Last updated on
Does it keep working when something breaks?
Reliability is the workload’s ability to deliver its promised function in the presence of failure. Not the absence of failure, which is not on offer. Disks fail, zones lose power, dependencies return errors, and someone deploys a bad configuration on a Friday afternoon. Reliability is what you have designed to happen next.
What this pillar covers
Section titled “What this pillar covers”- Deciding what “reliable enough” means for this workload, in business terms
- Identifying which flows matter and ranking them
- Understanding how the workload can fail, and deciding what happens when it does
- Redundancy, resilience patterns, and graceful degradation
- Backup, restore, and disaster recovery
- Proving all of the above through testing rather than assuming it
What it does not cover
Section titled “What it does not cover”Reliability against deliberate disruption belongs to Security, a denial-of- service attack and a traffic spike look similar and are addressed differently. The tooling and process for detecting and responding to incidents belongs to Operational Excellence; this pillar defines what needs detecting. Whether the system is fast enough belongs to Performance Efficiency, though the two meet at the point where saturation becomes an outage.
The central idea
Section titled “The central idea”Reliability is bought, not designed in for free. Every nine costs money, complexity, and operational load, and the cost is not linear. Which is why this pillar begins with a business conversation rather than a technical one: until you know what an hour of downtime costs and how much data loss the business can absorb, any architectural decision about redundancy is a guess dressed as engineering.
The second idea, which practitioners reach later than they expect: recovery matters more than prevention. Effort spent making failure less likely has a ceiling; effort spent making recovery fast, routine, and rehearsed does not. A workload that fails twice a year and recovers in ninety seconds is more reliable in every way the business cares about than one that fails once a year and takes six hours to come back.
Where to start
Section titled “Where to start”- Design principles: the reasoning, in five short statements
- Tradeoffs: what pursuing reliability costs the other pillars
Questions
Section titled “Questions”Ten questions. Numbers follow the order the decisions are usually made in and do not indicate priority.
Each links to its full page: the reasoning, what STACKIT provides, the tradeoffs, and the assessment questions.
REL 1: How do you derive reliability targets from business impact?
Section titled “REL 1: How do you derive reliability targets from business impact?”Establish availability targets, RTO, and RPO per flow, based on what an outage and what data loss actually cost the business. Record who agreed to them.
REL 2: How do you identify and rank the critical flows?
Section titled “REL 2: How do you identify and rank the critical flows?”Enumerate the paths through the workload that deliver business value, and rank them by consequence of failure. Reliability investment follows the ranking; not all components deserve equal treatment.
REL 3: How do you analyse failure modes and decide what happens when each occurs?
Section titled “REL 3: How do you analyse failure modes and decide what happens when each occurs?”For each component and dependency on a critical flow, identify how it can fail, what the effect is, and what the system does about it. An identified failure mode with no assigned mitigation is an accepted risk and must be recorded as one.
REL 4: How do you eliminate single points of failure?
Section titled “REL 4: How do you eliminate single points of failure?”Distribute components across availability zones, and across regions where the targets require it. Redundancy must cover state, not only compute, and the failover path itself must not become the single point of failure.
REL 5: How do you make remote interactions resilient?
Section titled “REL 5: How do you make remote interactions resilient?”Every call that crosses a process boundary needs a timeout and a retry policy with backoff and jitter; add a circuit breaker where a durably unhealthy dependency would otherwise consume the caller. Retried operations must be idempotent, or retries turn a transient fault into data corruption.
REL 6: How do you degrade gracefully and shed load deliberately?
Section titled “REL 6: How do you degrade gracefully and shed load deliberately?”Decide in advance which functions may be dropped when a dependency fails or demand exceeds capacity. Partial service is almost always better than total failure, and load shed on purpose is better than a system that collapses under its own queue.
REL 7: How do you scale to absorb demand, and how much headroom do you keep?
Section titled “REL 7: How do you scale to absorb demand, and how much headroom do you keep?”Size for realistic peaks, automate scaling where the workload allows it, and keep enough headroom to survive both the growth you predicted and the spike you did not. Know where the scaling limit is before you meet it.
REL 8: How do you back up data, and how do you know you can restore it?
Section titled “REL 8: How do you back up data, and how do you know you can restore it?”Back up every stateful component to a schedule derived from its RPO, hold copies where a single failure cannot destroy both, and restore from them on a regular cadence into a clean environment. An untested backup is not a backup.
REL 9: How do you plan and rehearse disaster recovery?
Section titled “REL 9: How do you plan and rehearse disaster recovery?”Document what happens when a whole region, a whole platform service, or the primary data set is lost: who decides, what the sequence is, what the dependencies are. Rehearse it, and time the rehearsal against the RTO you committed to.
REL 10: How do you know the workload is healthy, and how do you test that it stays so?
Section titled “REL 10: How do you know the workload is healthy, and how do you test that it stays so?”Define what “healthy” means per flow, in terms that map signals to a verdict rather than to a dashboard. Then verify continuously through fault injection, dependency failure simulation, and failover drills, in production where you can, in a production-like environment where you cannot.