Operational Excellence
Zuletzt aktualisiert am
Can a team run it without heroics?
Every other pillar describes a property of the workload. This one describes a property of the people and practices around it, which is why it is the pillar most often skipped in architecture reviews and the one whose absence eventually undermines all the others.
A workload with excellent reliability mechanisms that nobody knows how to operate is not reliable. A security baseline that no automated process enforces is not a baseline. A cost model nobody reviews is a document. Operational Excellence is the pillar that makes the rest of the framework hold over time rather than at the moment of design.
The test in the heading is deliberate. Not “can it be run”, anything can be run by a sufficiently dedicated engineer at three in the morning. The question is whether it can be run by a normal team on a normal day, and recovered by someone who did not build it.
What this pillar covers
Section titled “What this pillar covers”- Culture: shared ownership between the people who build and the people who run
- Development standards and automated quality gates
- Infrastructure, configuration, and policy as versioned code
- Deployment automation, and making deployments unremarkable
- Safe deployment practices: progressive exposure, health gates, rollback
- Environment consistency
- Observability, instrumentation, correlation, and signals that answer questions
- Operational procedures and runbooks
- Incident management and learning from failure
- Eliminating toil
What it does not cover
Section titled “What it does not cover”What needs to be reliable belongs to Reliability; this pillar covers the practices that keep those mechanisms working. What needs to be detected as a security event belongs to Security; the tooling and process for detection are shared.
Whether the system is fast enough, and keeping it fast as it and its traffic change, belongs to
Performance Efficiency. What is here is what every performance
question is answered from: the telemetry, in OPS 7, and the deployment gate that stops a
regression reaching everyone, in OPS 5.
The central idea
Section titled “The central idea”Everything that runs in production is defined as code, and everything routine is automated.
Not because automation is inherently virtuous, but because of what manual operations actually are: undocumented, unreviewed, unrepeatable, and unavailable when the person who knows how is asleep. A console click leaves no record of intent, cannot be reviewed before it happens, cannot be replicated in another environment, and cannot be rolled back. Every one of those properties matters most during an incident, which is exactly when manual operations are most likely.
The second idea, which sounds like a preference and is a design requirement: deployments should be boring. A deployment that is an event (scheduled, announced, requiring several people and a recovery plan) will happen rarely, batch up many changes, and be genuinely risky when it does. The rarity causes the risk, not the other way round. Deployments that are small, frequent, automated, and reversible are safer precisely because they are unremarkable, and a team that deploys daily recovers faster than one that deploys quarterly.
Where to start
Section titled “Where to start”Questions
Section titled “Questions”Ten questions. Numbers follow the order the decisions are usually made in and do not indicate priority.
OPS 1: How do you share operational responsibility between those who build and those who run?
Section titled “OPS 1: How do you share operational responsibility between those who build and those who run?”Give the people who build a system a stake in running it, and make incident review safe enough that the useful information actually surfaces. Both are structural decisions about how teams are organized, not statements of intent.
OPS 2: How do you define development standards and enforce them automatically?
Section titled “OPS 2: How do you define development standards and enforce them automatically?”Agree the standards (style, testing, review, dependency policy, branching) and enforce them in the pipeline rather than in review comments. A standard that depends on someone remembering is a suggestion.
OPS 3: How do you define infrastructure, configuration, and policy as versioned code?
Section titled “OPS 3: How do you define infrastructure, configuration, and policy as versioned code?”Everything that determines production behaviour lives in version control, is reviewed before it takes effect, and can be recreated from the repository. Detect drift and treat it as a defect.
OPS 4: How do you make deployment repeatable and reversible?
Section titled “OPS 4: How do you make deployment repeatable and reversible?”One automated path from source to production, used by everyone including for urgent fixes. Rollback must be a routine operation rather than an improvised one, which means it has to be exercised.
OPS 5: How do you limit the exposure of a bad change?
Section titled “OPS 5: How do you limit the exposure of a bad change?”Expose changes to a small fraction first, gate progression on health signals rather than elapsed time, and roll back automatically when the gate fails. A deployment strategy that depends on someone watching a dashboard does not work at three in the morning.
OPS 6: How do you keep environments consistent from development to production?
Section titled “OPS 6: How do you keep environments consistent from development to production?”Environments should differ in scale and data, not in shape. Divergence in topology, configuration mechanism, or platform version means testing in one tells you progressively less about the others.
OPS 7: How do you instrument the workload to answer questions you did not anticipate?
Section titled “OPS 7: How do you instrument the workload to answer questions you did not anticipate?”Emit signals that answer questions nobody thought to ask in advance, and propagate a correlation identifier through every hop so the three signal types can be joined into one narrative. Correlation must be designed in; it cannot be added at query time.
OPS 8: How do you document and rehearse operational procedures?
Section titled “OPS 8: How do you document and rehearse operational procedures?”Document the recurring operations and the responses to known failure modes, in enough detail that someone who did not build the system can execute them. Rehearse them: an unrehearsed runbook is a draft with unknown defects.
OPS 9: How do you manage incidents and learn from them?
Section titled “OPS 9: How do you manage incidents and learn from them?”Define severity, roles, communication, and escalation before you need them. Review every significant incident for structural causes rather than proximate ones, and track the resulting actions to completion: the review is worthless if its output is not.
OPS 10: How do you find and eliminate toil?
Section titled “OPS 10: How do you find and eliminate toil?”Track recurring manual operations, and treat them as defects with a cost rather than as the job. Toil scales with the system and consumes exactly the capacity that would have automated it away.
Related
Section titled “Related”- Design principles
- Tradeoffs
- Reliability:
OPS 8andREL 9overlap; rehearsing DR is where they meet