Skip to content
Beta

Operational Excellence

Last updated on

Can a team run it without heroics?

Every other pillar describes a property of the workload. This one describes a property of the people and practices around it, which is why it is the pillar most often skipped in architecture reviews and the one whose absence eventually undermines all the others.

A workload with excellent reliability mechanisms that nobody knows how to operate is not reliable. A security baseline that no automated process enforces is not a baseline. A cost model nobody reviews is a document. Operational Excellence is the pillar that makes the rest of the framework hold over time rather than at the moment of design.

The test in the heading is deliberate. Not “can it be run”, anything can be run by a sufficiently dedicated engineer at three in the morning. The question is whether it can be run by a normal team on a normal day, and recovered by someone who did not build it.

  • Culture: shared ownership between the people who build and the people who run
  • Development standards and automated quality gates
  • Infrastructure, configuration, and policy as versioned code
  • Deployment automation, and making deployments unremarkable
  • Safe deployment practices: progressive exposure, health gates, rollback
  • Environment consistency
  • Observability, instrumentation, correlation, and signals that answer questions
  • Operational procedures and runbooks
  • Incident management and learning from failure
  • Eliminating toil

What needs to be reliable belongs to Reliability; this pillar covers the practices that keep those mechanisms working. What needs to be detected as a security event belongs to Security; the tooling and process for detection are shared.

Whether the system is fast enough, and keeping it fast as it and its traffic change, belongs to Performance Efficiency. What is here is what every performance question is answered from: the telemetry, in OPS 7, and the deployment gate that stops a regression reaching everyone, in OPS 5.

Everything that runs in production is defined as code, and everything routine is automated.

Not because automation is inherently virtuous, but because of what manual operations actually are: undocumented, unreviewed, unrepeatable, and unavailable when the person who knows how is asleep. A console click leaves no record of intent, cannot be reviewed before it happens, cannot be replicated in another environment, and cannot be rolled back. Every one of those properties matters most during an incident, which is exactly when manual operations are most likely.

The second idea, which sounds like a preference and is a design requirement: deployments should be boring. A deployment that is an event (scheduled, announced, requiring several people and a recovery plan) will happen rarely, batch up many changes, and be genuinely risky when it does. The rarity causes the risk, not the other way round. Deployments that are small, frequent, automated, and reversible are safer precisely because they are unremarkable, and a team that deploys daily recovers faster than one that deploys quarterly.

  1. Design principles
  2. Tradeoffs

Ten questions. Numbers follow the order the decisions are usually made in and do not indicate priority.


OPS 1: How do you share operational responsibility between those who build and those who run?

Section titled “OPS 1: How do you share operational responsibility between those who build and those who run?”

Give the people who build a system a stake in running it, and make incident review safe enough that the useful information actually surfaces. Both are structural decisions about how teams are organized, not statements of intent.

→ Best practices

OPS 2: How do you define development standards and enforce them automatically?

Section titled “OPS 2: How do you define development standards and enforce them automatically?”

Agree the standards (style, testing, review, dependency policy, branching) and enforce them in the pipeline rather than in review comments. A standard that depends on someone remembering is a suggestion.

→ Best practices

OPS 3: How do you define infrastructure, configuration, and policy as versioned code?

Section titled “OPS 3: How do you define infrastructure, configuration, and policy as versioned code?”

Everything that determines production behaviour lives in version control, is reviewed before it takes effect, and can be recreated from the repository. Detect drift and treat it as a defect.

→ Best practices

One automated path from source to production, used by everyone including for urgent fixes. Rollback must be a routine operation rather than an improvised one, which means it has to be exercised.

→ Best practices

Expose changes to a small fraction first, gate progression on health signals rather than elapsed time, and roll back automatically when the gate fails. A deployment strategy that depends on someone watching a dashboard does not work at three in the morning.

→ Best practices

OPS 6: How do you keep environments consistent from development to production?

Section titled “OPS 6: How do you keep environments consistent from development to production?”

Environments should differ in scale and data, not in shape. Divergence in topology, configuration mechanism, or platform version means testing in one tells you progressively less about the others.

→ Best practices

OPS 7: How do you instrument the workload to answer questions you did not anticipate?

Section titled “OPS 7: How do you instrument the workload to answer questions you did not anticipate?”

Emit signals that answer questions nobody thought to ask in advance, and propagate a correlation identifier through every hop so the three signal types can be joined into one narrative. Correlation must be designed in; it cannot be added at query time.

→ Best practices

Document the recurring operations and the responses to known failure modes, in enough detail that someone who did not build the system can execute them. Rehearse them: an unrehearsed runbook is a draft with unknown defects.

→ Best practices

Define severity, roles, communication, and escalation before you need them. Review every significant incident for structural causes rather than proximate ones, and track the resulting actions to completion: the review is worthless if its output is not.

→ Best practices

Track recurring manual operations, and treat them as defects with a cost rather than as the job. Toil scales with the system and consumes exactly the capacity that would have automated it away.

→ Best practices