---
type: pillar
pillar: operational-excellence
code: OPS
title: Operational Excellence
description: Can a team run it without heroics? Everything as code, safe deployment, observability, incident practice and the elimination of toil, across ten questions.
status: draft
sidebar:
  label: "Overview"
  order: 0
hideLinkCard: true
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/"
source_file: "docs/architecture/pillars/operational-excellence/index.mdx"
---

> **Can a team run it without heroics?**

Every other pillar describes a property of the workload. This one describes a property of the people
and practices around it, which is why it is the pillar most often skipped in architecture reviews
and the one whose absence eventually undermines all the others.

A workload with excellent reliability mechanisms that nobody knows how to operate is not reliable.
A security baseline that no automated process enforces is not a baseline. A cost model nobody
reviews is a document. Operational Excellence is the pillar that makes the rest of the framework
hold over time rather than at the moment of design.

The test in the heading is deliberate. Not "can it be run", anything can be run by a sufficiently
dedicated engineer at three in the morning. The question is whether it can be run by a normal team
on a normal day, and recovered by someone who did not build it.

## What this pillar covers

- Culture: shared ownership between the people who build and the people who run
- Development standards and automated quality gates
- Infrastructure, configuration, and policy as versioned code
- Deployment automation, and making deployments unremarkable
- Safe deployment practices: progressive exposure, health gates, rollback
- Environment consistency
- Observability, instrumentation, correlation, and signals that answer questions
- Operational procedures and runbooks
- Incident management and learning from failure
- Eliminating toil

## What it does not cover

*What* needs to be reliable belongs to [Reliability](/architecture/pillars/reliability/); this pillar covers the
practices that keep those mechanisms working. *What* needs to be detected as a security event
belongs to [Security](/architecture/pillars/security/); the tooling and process for detection are shared.

Whether the system is fast enough, and keeping it fast as it and its traffic change, belongs to
[Performance Efficiency](/architecture/pillars/performance-efficiency/). What is here is what every performance
question is answered from: the telemetry, in [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/), and the deployment gate that stops a
regression reaching everyone, in [`OPS 5`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/).

## The central idea

**Everything that runs in production is defined as code, and everything routine is automated.**

Not because automation is inherently virtuous, but because of what manual operations actually are:
undocumented, unreviewed, unrepeatable, and unavailable when the person who knows how is asleep. A
console click leaves no record of intent, cannot be reviewed before it happens, cannot be
replicated in another environment, and cannot be rolled back. Every one of those properties
matters most during an incident, which is exactly when manual operations are most likely.

The second idea, which sounds like a preference and is a design requirement: **deployments should
be boring**. A deployment that is an event (scheduled, announced, requiring several people and a
recovery plan) will happen rarely, batch up many changes, and be genuinely risky when it does. The
rarity causes the risk, not the other way round. Deployments that are small, frequent, automated,
and reversible are safer precisely because they are unremarkable, and a team that deploys daily
recovers faster than one that deploys quarterly.

## Where to start

1. <LinkChip href="/architecture/pillars/operational-excellence/principles/">Design principles</LinkChip>
2. <LinkChip href="/architecture/pillars/operational-excellence/tradeoffs/">Tradeoffs</LinkChip>

## Questions

Ten questions. Numbers follow the order the decisions are usually made in and do **not** indicate
priority.

---

### [`OPS 1`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/): How do you share operational responsibility between those who build and those who run?

Give the people who build a system a stake in running it, and make incident review safe enough
that the useful information actually surfaces. Both are structural decisions about how teams are
organized, not statements of intent.

→ [Best practices](/architecture/pillars/operational-excellence/ops-01-shared-ownership/)

### [`OPS 2`](/architecture/pillars/operational-excellence/ops-02-development-standards/): How do you define development standards and enforce them automatically?

Agree the standards (style, testing, review, dependency policy, branching) and enforce them in the
pipeline rather than in review comments. A standard that depends on someone remembering is a
suggestion.

→ [Best practices](/architecture/pillars/operational-excellence/ops-02-development-standards/)

### [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/): How do you define infrastructure, configuration, and policy as versioned code?

Everything that determines production behaviour lives in version control, is reviewed before it
takes effect, and can be recreated from the repository. Detect drift and treat it as a defect.

→ [Best practices](/architecture/pillars/operational-excellence/ops-03-everything-as-code/)

### [`OPS 4`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/): How do you make deployment repeatable and reversible?

One automated path from source to production, used by everyone including for urgent fixes.
Rollback must be a routine operation rather than an improvised one, which means it has to be
exercised.

→ [Best practices](/architecture/pillars/operational-excellence/ops-04-deployment-automation/)

### [`OPS 5`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/): How do you limit the exposure of a bad change?

Expose changes to a small fraction first, gate progression on health signals rather than elapsed
time, and roll back automatically when the gate fails. A deployment strategy that depends on
someone watching a dashboard does not work at three in the morning.

→ [Best practices](/architecture/pillars/operational-excellence/ops-05-safe-deployment/)

### [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/): How do you keep environments consistent from development to production?

Environments should differ in scale and data, not in shape. Divergence in topology, configuration
mechanism, or platform version means testing in one tells you progressively less about the others.

→ [Best practices](/architecture/pillars/operational-excellence/ops-06-environment-consistency/)

### [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/): How do you instrument the workload to answer questions you did not anticipate?

Emit signals that answer questions nobody thought to ask in advance, and propagate a correlation
identifier through every hop so the three signal types can be joined into one narrative.
Correlation must be designed in; it cannot be added at query time.

→ [Best practices](/architecture/pillars/operational-excellence/ops-07-observability/)

### [`OPS 8`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/): How do you document and rehearse operational procedures?

Document the recurring operations and the responses to known failure modes, in enough detail that
someone who did not build the system can execute them. Rehearse them: an unrehearsed runbook is a
draft with unknown defects.

→ [Best practices](/architecture/pillars/operational-excellence/ops-08-operational-procedures/)

### [`OPS 9`](/architecture/pillars/operational-excellence/ops-09-incident-management/): How do you manage incidents and learn from them?

Define severity, roles, communication, and escalation before you need them. Review every
significant incident for structural causes rather than proximate ones, and track the resulting
actions to completion: the review is worthless if its output is not.

→ [Best practices](/architecture/pillars/operational-excellence/ops-09-incident-management/)

### [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/): How do you find and eliminate toil?

Track recurring manual operations, and treat them as defects with a cost rather than as the job.
Toil scales with the system and consumes exactly the capacity that would have automated it away.

→ [Best practices](/architecture/pillars/operational-excellence/ops-10-toil-elimination/)

## Related

- <LinkChip href="/architecture/pillars/operational-excellence/principles/">Design principles</LinkChip>
- <LinkChip href="/architecture/pillars/operational-excellence/tradeoffs/">Tradeoffs</LinkChip>
- [Reliability](/architecture/pillars/reliability/): [`OPS 8`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/) and [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/) overlap; rehearsing
  DR is where they meet
