Operational Excellence: design principles
Last updated on
1. Operations is a shared responsibility
Section titled “1. Operations is a shared responsibility”When the people who build a system are insulated from the consequences of operating it, the system acquires a predictable shape: convenient to write, unpleasant to run. Error messages that identify nothing. Configuration that requires tribal knowledge. Failure modes that produce a page and no diagnostic path. None of it is deliberate; it is what happens when the feedback loop is missing.
Closing the loop changes the design, not just the staffing. Engineers who carry a pager instrument their code differently, and they discover which of their failure modes are actually recoverable.
The corresponding half is blamelessness. A team that fears the consequences of an incident will optimize for not being implicated in one, fewer deployments, less experimentation, incident reports that omit the interesting part. The information you need to prevent the next failure is held by the person closest to this one, and it is only available if giving it is safe.
2. Everything that runs in production is defined as code
Section titled “2. Everything that runs in production is defined as code”Infrastructure, configuration, policy, pipelines, alerts, dashboards. Anything that determines how the system behaves belongs in version control.
The value is not the automation. It is the properties that come with it: a review before a change takes effect, a history explaining why a thing is the way it is, the ability to recreate an environment, the ability to detect drift, and the ability to revert.
A console click has none of those. It leaves no record of intent, cannot be reviewed in advance, cannot be reproduced elsewhere, and is discovered months later by someone asking why production does not match staging.
This is not absolutism about emergency access: sometimes the right thing is to fix production now. But an emergency change is followed by codifying it, or the drift is permanent.
3. Automate the routine; reserve humans for judgement
Section titled “3. Automate the routine; reserve humans for judgement”People are excellent at diagnosis, pattern recognition under ambiguity, and deciding what to do about a situation nobody anticipated. They are unreliable at executing a twelve-step procedure correctly at four in the morning for the fortieth time.
Give each what it is good at. Automate the procedures, the checks, the provisioning, the routine remediation, and use the freed attention for the problems that need thought.
Toil deserves naming because it disguises itself as work. Manual operations that recur feel productive, are visible, and generate gratitude. They also scale linearly with the system, consume the capacity that would have automated them away, and are where the errors come from. A recurring manual operation is a backlog item that has not been written down.
4. Deployments should be unremarkable
Section titled “4. Deployments should be unremarkable”The instinct that deployment is risky, and should therefore be rare, careful, and heavily ceremonied, is exactly backwards, and it is self-reinforcing.
Rare deployments accumulate changes. Large changes are harder to review, harder to test, and much harder to diagnose when something breaks, because the fault could be in any of forty things. Ceremony makes deployment expensive, which makes it rarer, which makes each one larger. The risk was created by the caution.
The inversion: small, frequent, automated, reversible. Each deployment carries one comprehensible change. A failure has an obvious suspect. Rollback is a routine operation rather than a decision requiring approval. Teams that deploy many times a day recover faster than teams that deploy quarterly, and they do so because the practice is exercised constantly rather than annually.
5. Observability is a design requirement, not a tooling choice
Section titled “5. Observability is a design requirement, not a tooling choice”You cannot operate what you cannot see, and you cannot see a system by installing something on top of it afterwards. Whether a system is observable is determined by what it emits, and that is a property of the code, decided when it is written.
The distinction that matters: monitoring tells you a known condition occurred. Observability lets you answer a question you had not thought to ask in advance. Dashboards for the failures you predicted are necessary and insufficient, because the incident that hurts is the one nobody predicted.
The practical requirement is correlation. Metrics, logs, and traces that cannot be joined across a single request produce three partial accounts and no narrative. The join key has to be designed in, propagated through every hop, because it cannot be added at query time.
6. Every incident is a lesson or a repeat
Section titled “6. Every incident is a lesson or a repeat”An incident is expensive information. Whether it was worth the price depends entirely on what happens afterwards.
The failure mode is a review that identifies a proximate cause, assigns a fix, and closes. Proximate causes are rarely interesting, the certificate expired, the disk filled, the configuration was wrong. The useful questions are structural: why did nothing catch it earlier, why was recovery slower than expected, what made this failure mode possible, and what else shares that property.
Which requires blamelessness to be real rather than stated. A review that concludes with a person’s name has found the cheapest available answer and stopped short of the useful one.
Related
Section titled “Related”- Overview: the questions this pillar asks
- Tradeoffs
- Reliability principles: reliability mechanisms are only as good as the practices that keep them working