OPS 8. How do you document and rehearse operational procedures?
Zuletzt aktualisiert am
Procedures exist so that nobody has to reason from first principles at three in the morning. That is their only purpose, and it sets the standard they have to meet: executable by a competent person who did not build the system, under pressure, without asking anyone.
Most runbooks fail that test. They were written by the author of the system, in the author’s vocabulary, and they have never been followed by anyone else, so their defects are unknown until the moment they matter.
Best practices
Section titled “Best practices”OPS 8.1Write a procedure for every recurring operation and known failure modeOPS 8.2Write for someone who did not build the systemOPS 8.3Rehearse them, and fix what the rehearsal findsOPS 8.4Keep procedures reachable when the environment is not
OPS 8.1 Write a procedure for every recurring operation and known failure mode
Section titled “OPS 8.1 Write a procedure for every recurring operation and known failure mode”Risk if not established: Medium
Two sources tell you what to write, and neither requires guessing.
The failure mode analysis from REL 3.2 recorded a response for each way a component can
fail. Every response that involves a human doing something is a procedure that does not yet exist.
The recurring operations are whatever people actually do: rotating a credential, scaling a
component, draining a node, clearing a stuck queue, onboarding a tenant. OPS 10.1 measures
these, and its list is the same list.
Not everything needs a document. A procedure executed monthly by a team that knows it well is a
candidate for automation instead, which is OPS 10.2. A procedure executed rarely, under
pressure, or by someone who does not know it well is where writing pays.
Prioritize by consequence and by rarity together. The procedure used twice a year on the most critical flow is the one where nobody remembers, and it is the one most likely to be missing.
On STACKIT. The content is yours; the platform contribution is that many procedures can be scripted rather than described. Run Command executes scripts and commands across virtual machines, with command templates for the recurring ones, which turns a documented sequence into an executed one.
A procedure that exists as a reviewed script has an advantage over prose: it cannot be ambiguous,
and it can be tested. That is the boundary between this question and OPS 10.
Tradeoffs. Operational Excellence, against itself. Every procedure is a document that ages with the system, and a stale procedure is worse than none because it carries authority. Write the ones that earn it.
Verify. List the failure modes from your REL 3 analysis whose response involves a human. How
many have a written procedure?
OPS 8.2 Write for someone who did not build the system
Section titled “OPS 8.2 Write for someone who did not build the system”Risk if not established: Medium
The person following a procedure during an incident is frequently not the person who wrote it, and may be several months removed from the context in which it was written. That is the audience.
Practically, that means naming things concretely rather than describing them. “Restart the service” assumes the reader knows which one and how to reach it; a named resource with the exact invocation does not. Include the commands, the resource names as they actually appear, and the expected output, so that the follower can tell whether a step worked.
Two things separate a usable procedure from a list of steps. Preconditions state what must be true before starting, including which permissions are needed, since discovering a missing permission halfway through is where recoveries stall. Verification states how to confirm the step worked, because a procedure without checkpoints fails silently and the follower continues.
State what to do when a step fails, at least for the steps where the answer is not obvious. A procedure that assumes every step succeeds is describing the easy case.
On STACKIT. Include the concrete access path. If a step requires the STACKIT Portal, the CLI or the API , say which and say which permissions it needs. The CLI usage examples are a good source for the exact invocation rather than approximating it.
Naming resources concretely is only possible when resource names are predictable, which is another
return on defining them in code under OPS 3.1.
Tradeoffs. Little beyond the writing effort. Precise procedures are longer than vague ones and that is the correct direction.
Verify. Hand your most important runbook to someone who did not write it. Can they execute it without asking a question? Where do they stop?
OPS 8.3 Rehearse them, and fix what the rehearsal finds
Section titled “OPS 8.3 Rehearse them, and fix what the rehearsal finds”Risk if not established: High
An unrehearsed procedure has unknown defects, and they cluster in the places nobody thought about: a permission that was never granted, a step that assumes a tool is installed, a resource name that changed six months ago, a command whose flags were deprecated.
Rehearse with the person who would actually do it, not with the author. The author knows the missing steps and will fill them in without noticing, which is exactly the failure the rehearsal exists to detect.
Set a cadence proportional to consequence and inversely to frequency. A procedure executed weekly rehearses itself. One executed once a year on a critical flow needs a deliberate exercise, and that is also the one people assume is fine.
The findings are the output. A rehearsal that produced no corrections either rehearsed a procedure that is genuinely current or was performed by someone who knew the gaps. Both are worth distinguishing.
On STACKIT. Rehearsing needs somewhere to rehearse, which is OPS 6: an environment close
enough in shape that the procedure exercises the same behaviour. Creating one on demand from the
definitions in OPS 3 is what makes this affordable.
Where a procedure is scripted through Run Command , the rehearsal is also a test of the script, which is cheaper to run and gives a clearer pass or fail than following prose.
Tradeoffs. Cost Optimization. Rehearsal time produces nothing visible, and the environment
it needs costs something while it exists. It is the same argument as REL 8.3 and REL 9.3,
which are the two rehearsals with the highest return.
Verify. When was each critical procedure last rehearsed, by whom, and what did the rehearsal change? If nothing changed, was the rehearser the author?
OPS 8.4 Keep procedures reachable when the environment is not
Section titled “OPS 8.4 Keep procedures reachable when the environment is not”Risk if not established: Medium
A runbook stored only in the environment it describes is unavailable exactly when it is needed. The same applies to whatever the procedure depends on: credentials, contact details, artefacts, and the tools it assumes are installed.
This is the same circular dependency problem as REL 9.4, applied at a smaller scale and more
often. It is worth checking per procedure rather than once, because the dependencies differ.
Keep procedures next to the code they describe so they are versioned and reviewed with it, and keep a readable copy somewhere with an independent failure domain. Those two goals conflict slightly and both matter; a periodic export is usually enough to satisfy the second.
Test the access path rather than assuming it. The question is not whether a copy exists but whether the person on call at three in the morning can reach it with the credentials they have at that moment.
On STACKIT. Procedures belong in STACKIT Git alongside the code, for the versioning and review that gives. That also puts them inside a failure domain, so the independent copy is the complement rather than the alternative.
Authentication to the platform is itself on the path: if executing the procedure requires the
Portal, the CLI or the API, then identity is a dependency, which is the point REL 3.3 makes
generally and SOV 9 addresses from the sovereignty angle.
Tradeoffs. Security. Copies of procedures outside the primary environment are additional places that information can leak from, and procedures frequently name resources and access paths. They need protection proportional to what they reveal.
Verify. If your primary environment were unavailable, could the person on call read the relevant runbook and authenticate to act on it? When was that last tested rather than assumed?
Related
Section titled “Related”REL 3.2Failure mode responses, which supply what to write procedures forREL 9Disaster recovery, the largest procedure and the one most needing rehearsalOPS 6Environment consistency, which makes rehearsal meaningfulOPS 10Toil elimination, which is where a well-rehearsed procedure often ends upOPS 3.1Everything as code, which makes resource names predictable enough to write down