---
id: OPS08
pillar: operational-excellence
title: OPS 8. How do you document and rehearse operational procedures?
description: An unrehearsed runbook is a draft with unknown defects. How to write procedures for someone who did not build the system, and find the gaps before an incident.
status: draft
services: [automation-service]
sidebar:
  order: 17
  label: Operational procedures
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/ops-08-operational-procedures/"
source_file: "docs/architecture/pillars/operational-excellence/ops-08-operational-procedures.mdx"
---

Procedures exist so that nobody has to reason from first principles at three in the morning. That
is their only purpose, and it sets the standard they have to meet: executable by a competent
person who did not build the system, under pressure, without asking anyone.

Most runbooks fail that test. They were written by the author of the system, in the author's
vocabulary, and they have never been followed by anyone else, so their defects are unknown until
the moment they matter.

## Best practices

- [`OPS 8.1`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/#ops-81-write-a-procedure-for-every-recurring-operation-and-known-failure-mode) Write a procedure for every recurring operation and known failure mode
- [`OPS 8.2`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/#ops-82-write-for-someone-who-did-not-build-the-system) Write for someone who did not build the system
- [`OPS 8.3`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/#ops-83-rehearse-them-and-fix-what-the-rehearsal-finds) Rehearse them, and fix what the rehearsal finds
- [`OPS 8.4`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/#ops-84-keep-procedures-reachable-when-the-environment-is-not) Keep procedures reachable when the environment is not

---

## OPS 8.1 Write a procedure for every recurring operation and known failure mode

**Risk if not established:** Medium

Two sources tell you what to write, and neither requires guessing.

The **failure mode analysis** from [`REL 3.2`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-32-decide-and-record-the-response-to-each-failure-mode) recorded a response for each way a component can
fail. Every response that involves a human doing something is a procedure that does not yet exist.

The **recurring operations** are whatever people actually do: rotating a credential, scaling a
component, draining a node, clearing a stuck queue, onboarding a tenant. [`OPS 10.1`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-101-measure-recurring-manual-work-rather-than-absorbing-it) measures
these, and its list is the same list.

Not everything needs a document. A procedure executed monthly by a team that knows it well is a
candidate for automation instead, which is [`OPS 10.2`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-102-automate-what-recurs-starting-with-the-riskiest-rather-than-the-most-frequent). A procedure executed rarely, under
pressure, or by someone who does not know it well is where writing pays.

Prioritize by consequence and by rarity together. The procedure used twice a year on the most
critical flow is the one where nobody remembers, and it is the one most likely to be missing.

**On STACKIT.** The content is yours; the platform contribution is that many procedures can be
scripted rather than described. <LinkChip href="https://docs.stackit.cloud/products/compute-engine/run-command/">Run
Command</LinkChip> executes scripts and
commands across virtual machines, with <LinkChip href="https://docs.stackit.cloud/products/compute-engine/run-command/getting-started/command-templates/">command
templates</LinkChip>
for the recurring ones, which turns a documented sequence into an executed one.

A procedure that exists as a reviewed script has an advantage over prose: it cannot be ambiguous,
and it can be tested. That is the boundary between this question and [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/).

**Tradeoffs.** **Operational Excellence**, against itself. Every procedure is a document that ages
with the system, and a stale procedure is worse than none because it carries authority. Write the
ones that earn it.

**Verify.** List the failure modes from your [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) analysis whose response involves a human. How
many have a written procedure?

---

## OPS 8.2 Write for someone who did not build the system

**Risk if not established:** Medium

The person following a procedure during an incident is frequently not the person who wrote it, and
may be several months removed from the context in which it was written. That is the audience.

Practically, that means naming things concretely rather than describing them. "Restart the
service" assumes the reader knows which one and how to reach it; a named resource with the exact
invocation does not. Include the commands, the resource names as they actually appear, and the
expected output, so that the follower can tell whether a step worked.

Two things separate a usable procedure from a list of steps. **Preconditions** state what must be
true before starting, including which permissions are needed, since discovering a missing
permission halfway through is where recoveries stall. **Verification** states how to confirm the
step worked, because a procedure without checkpoints fails silently and the follower continues.

State what to do when a step fails, at least for the steps where the answer is not obvious. A
procedure that assumes every step succeeds is describing the easy case.

**On STACKIT.** Include the concrete access path. If a step requires the STACKIT Portal, the
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-cli/">CLI</LinkChip> or the
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-api/">API</LinkChip>, say which and say which
permissions it needs. The <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-cli/usage-examples/">CLI usage
examples</LinkChip> are a good
source for the exact invocation rather than approximating it.

Naming resources concretely is only possible when resource names are predictable, which is another
return on defining them in code under [`OPS 3.1`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-31-define-every-production-resource-in-version-control).

**Tradeoffs.** Little beyond the writing effort. Precise procedures are longer than vague ones and
that is the correct direction.

**Verify.** Hand your most important runbook to someone who did not write it. Can they execute it
without asking a question? Where do they stop?

---

## OPS 8.3 Rehearse them, and fix what the rehearsal finds

**Risk if not established:** High

An unrehearsed procedure has unknown defects, and they cluster in the places nobody thought about:
a permission that was never granted, a step that assumes a tool is installed, a resource name that
changed six months ago, a command whose flags were deprecated.

Rehearse with the person who would actually do it, not with the author. The author knows the
missing steps and will fill them in without noticing, which is exactly the failure the rehearsal
exists to detect.

Set a cadence proportional to consequence and inversely to frequency. A procedure executed weekly
rehearses itself. One executed once a year on a critical flow needs a deliberate exercise, and
that is also the one people assume is fine.

The findings are the output. A rehearsal that produced no corrections either rehearsed a procedure
that is genuinely current or was performed by someone who knew the gaps. Both are worth
distinguishing.

**On STACKIT.** Rehearsing needs somewhere to rehearse, which is [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/): an environment close
enough in shape that the procedure exercises the same behaviour. Creating one on demand from the
definitions in [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) is what makes this affordable.

Where a procedure is scripted through
<LinkChip href="https://docs.stackit.cloud/products/compute-engine/run-command/">Run Command</LinkChip>, the rehearsal is
also a test of the script, which is cheaper to run and gives a clearer pass or fail than following
prose.

**Tradeoffs.** **Cost Optimization.** Rehearsal time produces nothing visible, and the environment
it needs costs something while it exists. It is the same argument as [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it) and [`REL 9.3`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-93-rehearse-it-and-time-the-rehearsal-against-the-rto),
which are the two rehearsals with the highest return.

**Verify.** When was each critical procedure last rehearsed, by whom, and what did the rehearsal
change? If nothing changed, was the rehearser the author?

---

## OPS 8.4 Keep procedures reachable when the environment is not

**Risk if not established:** Medium

A runbook stored only in the environment it describes is unavailable exactly when it is needed.
The same applies to whatever the procedure depends on: credentials, contact details, artefacts,
and the tools it assumes are installed.

This is the same circular dependency problem as [`REL 9.4`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-94-keep-the-plan-and-its-dependencies-reachable-when-the-primary-environment-is-not), applied at a smaller scale and more
often. It is worth checking per procedure rather than once, because the dependencies differ.

Keep procedures next to the code they describe so they are versioned and reviewed with it, and
keep a readable copy somewhere with an independent failure domain. Those two goals conflict
slightly and both matter; a periodic export is usually enough to satisfy the second.

Test the access path rather than assuming it. The question is not whether a copy exists but
whether the person on call at three in the morning can reach it with the credentials they have at
that moment.

**On STACKIT.** Procedures belong in
<LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/">STACKIT Git</LinkChip> alongside the code, for
the versioning and review that gives. That also puts them inside a failure domain, so the
independent copy is the complement rather than the alternative.

Authentication to the platform is itself on the path: if executing the procedure requires the
Portal, the CLI or the API, then identity is a dependency, which is the point [`REL 3.3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-33-include-the-failure-modes-of-your-dependencies-not-only-of-your-own-code) makes
generally and [`SOV 9`](/architecture/pillars/sovereignty/sov-09-identity-sovereignty/) addresses from the sovereignty angle.

**Tradeoffs.** **Security.** Copies of procedures outside the primary environment are additional
places that information can leak from, and procedures frequently name resources and access paths.
They need protection proportional to what they reveal.

**Verify.** If your primary environment were unavailable, could the person on call read the
relevant runbook and authenticate to act on it? When was that last tested rather than assumed?

---

## Related

- [`REL 3.2`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-32-decide-and-record-the-response-to-each-failure-mode) Failure mode responses, which supply what to write procedures for
- [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/) Disaster recovery, the largest procedure and the one most needing rehearsal
- [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) Environment consistency, which makes rehearsal meaningful
- [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/) Toil elimination, which is where a well-rehearsed procedure often ends up
- [`OPS 3.1`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-31-define-every-production-resource-in-version-control) Everything as code, which makes resource names predictable enough to write down
