---
type: principles
pillar: reliability
code: REL
title: "Reliability: design principles"
description: "Five principles behind the Reliability pillar: why reliability is a business decision and why recovery beats prevention."
status: draft
sidebar:
  label: "Design principles"
  order: 1
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/principles/"
source_file: "docs/architecture/pillars/reliability/principles.mdx"
---

Principles are not checkable and never appear in an assessment. They are here so that you can tell
when a best practice does not apply to your situation, which is a judgement no list of questions can make
for you.

---

## 1. Reliability is a business decision

How reliable a workload should be is not an engineering question. It is a question about what
downtime costs, what data loss costs, and what the organization is willing to pay to avoid them.
Engineering answers *how*, once someone has answered *how much*.

This gets skipped constantly, usually because the business conversation is harder than the
technical one. The result is a workload that is over-engineered in the places the team found
interesting and under-engineered everywhere else, with no way to tell which is which.

The corollary is uncomfortable but freeing: for some workloads the right answer is *less*
reliability than you are currently building. An internal reporting tool that nobody looks at
between 6pm and 8am does not need multi-zone redundancy, and the money is better spent elsewhere.

## 2. Failure is normal: design for it, not against it

Every component you depend on will eventually fail, and the ones you did not think of as
dependencies will fail too. A design that assumes healthy dependencies is not a design; it is an
optimistic sketch.

The practical shift is to stop asking "how do I prevent this?" and start asking "what happens when
this fails, and is that acceptable?" The first question has diminishing returns and no natural
stopping point. The second one has an answer you can verify.

## 3. Simplicity is a reliability feature

Every component is a thing that can fail, needs patching, holds a misconfiguration, and has to be
reasoned about at three in the morning by someone who did not build it. Complexity added in the
name of reliability regularly costs more reliability than it buys.

The failure mode is specific and common: a redundancy mechanism that is itself a single point of
failure, or a failover path so intricate that it has never been successfully exercised. If a
resilience mechanism cannot be explained to a new team member in a few minutes, it will not work
during an incident.

Prefer the boring option. Fewer components, fewer states, fewer conditional paths. Reach for a
managed service over an assembly of parts you maintain yourself, not because managed services do
not fail, but because their failure modes are documented and someone else is paid to fix them.

## 4. Recovery beats prevention

Preventing failure has a ceiling. Recovering from it does not.

Optimize for the time between "something broke" and "customers stopped noticing". That means fast
detection, small blast radius, rehearsed procedures, and the ability to roll back without a
meeting. Most organizations discover this in the wrong order, they spend years hardening against
failures that keep happening anyway, then discover that halving their recovery time was cheaper
and helped with every failure including the ones they never predicted.

## 5. Untested reliability is assumed reliability

A backup that has never been restored is not a backup. A failover that has never been triggered is
a theory. A runbook nobody has followed is a document.

Reliability mechanisms decay silently: the backup job that has been failing for six weeks, the
standby whose configuration drifted, the certificate on the failover path that expired. None of
these announce themselves. They are discovered during the incident, which is the most expensive
possible time to discover them.

The only reliable signal is regular, deliberate exercise, restoring into a clean environment,
failing over on purpose, injecting the fault and watching what the system does. If that sounds
risky, then that is precisely the information you were looking for.

---

## Related

- [Overview](/architecture/pillars/reliability/): the questions this pillar asks
- <LinkChip href="/architecture/pillars/reliability/tradeoffs/">Tradeoffs</LinkChip>
- [Operational Excellence principles](/architecture/pillars/operational-excellence/principles/): reliability
  mechanisms are only as good as the practices that operate them
