Skip to content
Beta

Reliability: design principles

Last updated on

Principles are not checkable and never appear in an assessment. They are here so that you can tell when a best practice does not apply to your situation, which is a judgement no list of questions can make for you.


How reliable a workload should be is not an engineering question. It is a question about what downtime costs, what data loss costs, and what the organization is willing to pay to avoid them. Engineering answers how, once someone has answered how much.

This gets skipped constantly, usually because the business conversation is harder than the technical one. The result is a workload that is over-engineered in the places the team found interesting and under-engineered everywhere else, with no way to tell which is which.

The corollary is uncomfortable but freeing: for some workloads the right answer is less reliability than you are currently building. An internal reporting tool that nobody looks at between 6pm and 8am does not need multi-zone redundancy, and the money is better spent elsewhere.

2. Failure is normal: design for it, not against it

Section titled “2. Failure is normal: design for it, not against it”

Every component you depend on will eventually fail, and the ones you did not think of as dependencies will fail too. A design that assumes healthy dependencies is not a design; it is an optimistic sketch.

The practical shift is to stop asking “how do I prevent this?” and start asking “what happens when this fails, and is that acceptable?” The first question has diminishing returns and no natural stopping point. The second one has an answer you can verify.

Every component is a thing that can fail, needs patching, holds a misconfiguration, and has to be reasoned about at three in the morning by someone who did not build it. Complexity added in the name of reliability regularly costs more reliability than it buys.

The failure mode is specific and common: a redundancy mechanism that is itself a single point of failure, or a failover path so intricate that it has never been successfully exercised. If a resilience mechanism cannot be explained to a new team member in a few minutes, it will not work during an incident.

Prefer the boring option. Fewer components, fewer states, fewer conditional paths. Reach for a managed service over an assembly of parts you maintain yourself, not because managed services do not fail, but because their failure modes are documented and someone else is paid to fix them.

Preventing failure has a ceiling. Recovering from it does not.

Optimize for the time between “something broke” and “customers stopped noticing”. That means fast detection, small blast radius, rehearsed procedures, and the ability to roll back without a meeting. Most organizations discover this in the wrong order, they spend years hardening against failures that keep happening anyway, then discover that halving their recovery time was cheaper and helped with every failure including the ones they never predicted.

5. Untested reliability is assumed reliability

Section titled “5. Untested reliability is assumed reliability”

A backup that has never been restored is not a backup. A failover that has never been triggered is a theory. A runbook nobody has followed is a document.

Reliability mechanisms decay silently: the backup job that has been failing for six weeks, the standby whose configuration drifted, the certificate on the failover path that expired. None of these announce themselves. They are discovered during the incident, which is the most expensive possible time to discover them.

The only reliable signal is regular, deliberate exercise, restoring into a clean environment, failing over on purpose, injecting the fault and watching what the system does. If that sounds risky, then that is precisely the information you were looking for.