---
id: OPS03
pillar: operational-excellence
title: OPS 3. How do you define infrastructure, configuration, and policy as versioned code?
description: A console click leaves no record of intent, cannot be reviewed, and cannot be reproduced. What belongs in version control, and how to spot when reality drifts.
status: draft
services: [git]
sidebar:
  order: 12
  label: Everything as code
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/ops-03-everything-as-code/"
source_file: "docs/architecture/pillars/operational-excellence/ops-03-everything-as-code.mdx"
---

This is the question with the widest return in the pillar. It is a prerequisite for environment
consistency in [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/), for the recovery sequence in [`REL 9.2`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-92-write-the-sequence-including-who-decides-to-invoke-it), for the rehearsals in [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it),
and for most of what makes [`SUS 6`](/architecture/pillars/sustainability/sus-06-shut-down-idle/) executable.

The value is not the automation. It is the four properties that come with it: a change can be
reviewed before it takes effect, the history explains why something is the way it is, an
environment can be recreated, and divergence can be detected. A console click has none of those.

## Best practices

- [`OPS 3.1`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-31-define-every-production-resource-in-version-control) Define every production resource in version control
- [`OPS 3.2`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-32-review-changes-before-they-take-effect) Review changes before they take effect
- [`OPS 3.3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-33-detect-drift-and-treat-it-as-a-defect) Detect drift and treat it as a defect
- [`OPS 3.4`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-34-codify-emergency-changes-afterwards) Codify emergency changes afterwards

---

## OPS 3.1 Define every production resource in version control

**Risk if not established:** High

The scope is wider than most teams start with. Compute and networking are the obvious part.
The parts that get left behind are where recovery fails:

- **Access and roles.** Who can do what, which is also the audit evidence [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) needs.
- **Monitoring and alerting.** Dashboards and alert rules configured by hand are lost with the
  environment they lived in.
- **Policy.** Network rules, admission controllers, quota assignments.
- **The pipeline itself.** A build system configured through a web interface is a single point of
  failure with no history.
- **DNS records and certificates.** Small, critical, and almost never captured.

Secrets are the deliberate exception. Their existence and their consumers belong in code; their
values belong in a secret store, which is [`SEC 9`](/architecture/pillars/security/sec-09-secrets/).

Start with what has to be recreated during recovery, since that is the part with a demonstrable
cost of omission, and expand from there.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-iac/stackit-terraform-provider/">Terraform
provider</LinkChip>
covers the platform resources and is the usual starting point. There is also a <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-iac/pulumi/">Pulumi
provider</LinkChip> for teams that prefer
general-purpose languages.

Underneath both, the <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-api/">STACKIT API</LinkChip> and
the <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-cli/">CLI</LinkChip> are available, and there are
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-sdk/stackit-go-sdk/">SDKs for Go, Java and
Python</LinkChip> where automation
needs to be embedded in an application. Anything reachable through the API can be codified.

The <LinkChip href="https://docs.stackit.cloud/platform/resource-manager/">Resource Manager</LinkChip>
hierarchy is worth defining as code
too, not only the resources inside it. The project structure is what [`OPS 1.3`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-13-define-ownership-so-that-every-component-has-a-name-against-it), [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) and
[`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) all rest on, and a hierarchy that grew by hand is the one nobody can reproduce.

The provider covers most of the platform, and where it lags it lags predictably. A new feature
updates the API specification, the Go SDK regenerates from it automatically, and the Terraform
resource is then written by hand. That last step is the one that takes time, so the gaps sit at
the newest features rather than being scattered at random, and they close.

Two things follow for a design. Check the provider for the specific resources your workload needs
rather than assuming coverage, and do it before the recovery procedure in [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/) depends on it: a
resource that has to be created by hand is a manual step in a sequence you meant to automate, and
the worst time to discover it is while executing that sequence. Where a gap exists, the provider's
<LinkChip href="https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs/ephemeral-resources/access_token">ephemeral access
token</LinkChip>
lets Terraform authenticate against the API directly, so the gap can be bridged inside the same
run instead of beside it.

**Tradeoffs.** **Cost Optimization.** Real up-front engineering time, and a learning curve for
teams new to it. The return is in [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/), [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it) and [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/), all of which are difficult to
the point of impractical without it.

**Verify.** Could you recreate your production environment from the repository alone? List what
would be missing, and how you would know.

---

## OPS 3.2 Review changes before they take effect

**Risk if not established:** High

Reviewing infrastructure changes is the single largest security and reliability return in this
question, and it is available only because the changes are code.

The mechanism that makes it work is showing the effect rather than the intent. A plan output that
says a database will be replaced is information a reviewer can act on; a diff of the source
frequently is not, because the consequence of a parameter change is not always visible from the
parameter.

Apply the same discipline as for application code: someone other than the author, before it
reaches production, with the pipeline enforcing it rather than convention. [`OPS 2.2`](/architecture/pillars/operational-excellence/ops-02-development-standards/#ops-22-enforce-in-the-pipeline-rather-than-in-review-comments) is what makes
that reliable.

Give destructive changes their own treatment. Resource replacement and deletion deserve to be
visible in the review rather than buried, because they are where the expensive mistakes are.

**On STACKIT.** A plan step in <LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/basics/stackit-pipelines/">STACKIT
Pipelines</LinkChip>
posted for review before an apply step is the usual shape, and the platform supports <LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/how-tos/pipeline-manual-approval/">manual
approval in a
workflow</LinkChip>,
which is the gate that separates plan from apply.

The identity the pipeline uses is worth attention: it is frequently more privileged than any human
account, and it belongs under the same scrutiny, which is [`SEC 4`](/architecture/pillars/security/sec-04-identity/) and [`SEC 5`](/architecture/pillars/security/sec-05-least-privilege/). Actions taken with
it appear in the <LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> like any other.

**Tradeoffs.** **Operational Excellence**, against itself: review adds latency to every change,
which pushes against the small and frequent changes [`OPS 4.3`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-43-keep-changes-small-and-deploy-frequently) wants. Resolve it by making review
fast rather than optional, and by scoping approval requirements to changes that warrant them.

**Verify.** For your last ten infrastructure changes, how many were reviewed before they took
effect, and did the reviewer see the planned effect or only the source diff?

---

## OPS 3.3 Detect drift and treat it as a defect

**Risk if not established:** Medium

Drift is the gap between what the repository says and what exists. It appears through emergency
changes, console edits, external automation and manual experiments that were never cleaned up.

Undetected drift removes the value of everything else in this question. The repository stops
describing reality, recreating an environment produces something different from what was running,
and the review process governs a fiction.

Detect it by comparing periodically and automatically, not by trusting that nobody clicked. Then
decide per instance: either the change was wrong and reality should be corrected, or it was right
and the code should be updated. Both are legitimate; leaving it is not.

Distinguish drift from expected variance. Autoscaled node counts and generated identifiers change
without anyone editing anything, and a detector that reports them trains people to ignore the
report.

**On STACKIT.** Drift detection is a property of your tooling rather than of the platform. A
scheduled plan run in the pipeline that fails when it finds unexpected changes is the simplest
version and usually sufficient.

The <LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> is the complementary source: it
records every action by users, service accounts and the platform, at organization, folder and
project scope, and is enabled by default. When drift is found, that is where you look for what
changed and who changed it. Note the **90-day retention** in the Portal, which bounds how far back
an investigation can reach unless the records are exported. See [`OPS 7.4`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-74-set-retention-by-the-value-of-each-signal-rather-than-uniformly).

**Tradeoffs.** **Operational Excellence.** A drift detector is a thing to operate and tune, and a
noisy one is worse than none. **Cost Optimization.** Scheduled plan runs consume pipeline
capacity.

**Verify.** When was drift last detected in your production environment, how was it found, and
what was done about it? If the answer is that it has never been detected, is that because there is
none?

---

## OPS 3.4 Codify emergency changes afterwards

**Risk if not established:** Medium

Sometimes the right thing is to fix production now. A rule that forbids it produces either a
slower incident response or a rule everyone ignores, and the second is worse because it also
removes the record.

Permit the emergency change and require the follow-up. The change is made, the incident ends, and
then the same change is expressed in code so that the repository and reality agree again. Without
that second step, the drift is permanent and it accumulates precisely in the components that have
the most incidents.

Make the follow-up a tracked item with an owner and a deadline, in the same way incident actions
are handled in [`OPS 9.4`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-94-track-review-actions-to-completion). An intention to codify it later has the same completion rate as any
other unowned intention.

Watch the pattern rather than the individual instance. A component that repeatedly needs emergency
changes is telling you something about its design or its automation, which is [`OPS 10.3`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-103-remove-the-need-rather-than-automating-the-workaround).

**On STACKIT.** Emergency access through the STACKIT Portal or the CLI is recorded in the <LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit
log</LinkChip> by default, which gives the follow-up a
factual basis: what was actually changed, by whom, and when. That is more reliable than
reconstructing it from memory after a long night.

Where emergency access requires elevated permissions, the break-glass treatment in [`SEC 5`](/architecture/pillars/security/sec-05-least-privilege/)
applies: tightly bounded, heavily logged, and reviewed after every use.

**Tradeoffs.** **Security.** An emergency path that bypasses review is a real exposure, which is
why it is bounded and logged rather than removed. The tension is named in the [Security
tradeoffs](/architecture/pillars/security/tradeoffs/).

**Verify.** For the last emergency change made outside the normal process, was it subsequently
expressed in code, and how long did that take?

---

## Related

- [`OPS 2`](/architecture/pillars/operational-excellence/ops-02-development-standards/) Development standards, which the pipeline also enforces
- [`OPS 4`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/) Deployment automation, which runs on the same definitions
- [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) Environment consistency, which is impossible without this
- [`REL 9.2`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-92-write-the-sequence-including-who-decides-to-invoke-it) Disaster recovery, whose sequence is executable only if the environment is codified
- [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) Segmentation and [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) Cost attribution, which share the resource hierarchy
