Zum Inhalt springen
Beta

OPS 3. How do you define infrastructure, configuration, and policy as versioned code?

Zuletzt aktualisiert am

This is the question with the widest return in the pillar. It is a prerequisite for environment consistency in OPS 6, for the recovery sequence in REL 9.2, for the rehearsals in REL 8.3, and for most of what makes SUS 6 executable.

The value is not the automation. It is the four properties that come with it: a change can be reviewed before it takes effect, the history explains why something is the way it is, an environment can be recreated, and divergence can be detected. A console click has none of those.

  • OPS 3.1 Define every production resource in version control
  • OPS 3.2 Review changes before they take effect
  • OPS 3.3 Detect drift and treat it as a defect
  • OPS 3.4 Codify emergency changes afterwards

OPS 3.1 Define every production resource in version control

Section titled “OPS 3.1 Define every production resource in version control”

Risk if not established: High

The scope is wider than most teams start with. Compute and networking are the obvious part. The parts that get left behind are where recovery fails:

  • Access and roles. Who can do what, which is also the audit evidence SOV 7 needs.
  • Monitoring and alerting. Dashboards and alert rules configured by hand are lost with the environment they lived in.
  • Policy. Network rules, admission controllers, quota assignments.
  • The pipeline itself. A build system configured through a web interface is a single point of failure with no history.
  • DNS records and certificates. Small, critical, and almost never captured.

Secrets are the deliberate exception. Their existence and their consumers belong in code; their values belong in a secret store, which is SEC 9.

Start with what has to be recreated during recovery, since that is the part with a demonstrable cost of omission, and expand from there.

On STACKIT. The Terraform provider covers the platform resources and is the usual starting point. There is also a Pulumi provider for teams that prefer general-purpose languages.

Underneath both, the STACKIT API and the CLI are available, and there are SDKs for Go, Java and Python where automation needs to be embedded in an application. Anything reachable through the API can be codified.

The Resource Manager hierarchy is worth defining as code too, not only the resources inside it. The project structure is what OPS 1.3, SEC 2 and COST 2 all rest on, and a hierarchy that grew by hand is the one nobody can reproduce.

The provider covers most of the platform, and where it lags it lags predictably. A new feature updates the API specification, the Go SDK regenerates from it automatically, and the Terraform resource is then written by hand. That last step is the one that takes time, so the gaps sit at the newest features rather than being scattered at random, and they close.

Two things follow for a design. Check the provider for the specific resources your workload needs rather than assuming coverage, and do it before the recovery procedure in REL 9 depends on it: a resource that has to be created by hand is a manual step in a sequence you meant to automate, and the worst time to discover it is while executing that sequence. Where a gap exists, the provider’s ephemeral access token lets Terraform authenticate against the API directly, so the gap can be bridged inside the same run instead of beside it.

Tradeoffs. Cost Optimization. Real up-front engineering time, and a learning curve for teams new to it. The return is in OPS 6, REL 8.3 and REL 9, all of which are difficult to the point of impractical without it.

Verify. Could you recreate your production environment from the repository alone? List what would be missing, and how you would know.


OPS 3.2 Review changes before they take effect

Section titled “OPS 3.2 Review changes before they take effect”

Risk if not established: High

Reviewing infrastructure changes is the single largest security and reliability return in this question, and it is available only because the changes are code.

The mechanism that makes it work is showing the effect rather than the intent. A plan output that says a database will be replaced is information a reviewer can act on; a diff of the source frequently is not, because the consequence of a parameter change is not always visible from the parameter.

Apply the same discipline as for application code: someone other than the author, before it reaches production, with the pipeline enforcing it rather than convention. OPS 2.2 is what makes that reliable.

Give destructive changes their own treatment. Resource replacement and deletion deserve to be visible in the review rather than buried, because they are where the expensive mistakes are.

On STACKIT. A plan step in STACKIT Pipelines posted for review before an apply step is the usual shape, and the platform supports manual approval in a workflow , which is the gate that separates plan from apply.

The identity the pipeline uses is worth attention: it is frequently more privileged than any human account, and it belongs under the same scrutiny, which is SEC 4 and SEC 5. Actions taken with it appear in the audit log like any other.

Tradeoffs. Operational Excellence, against itself: review adds latency to every change, which pushes against the small and frequent changes OPS 4.3 wants. Resolve it by making review fast rather than optional, and by scoping approval requirements to changes that warrant them.

Verify. For your last ten infrastructure changes, how many were reviewed before they took effect, and did the reviewer see the planned effect or only the source diff?


OPS 3.3 Detect drift and treat it as a defect

Section titled “OPS 3.3 Detect drift and treat it as a defect”

Risk if not established: Medium

Drift is the gap between what the repository says and what exists. It appears through emergency changes, console edits, external automation and manual experiments that were never cleaned up.

Undetected drift removes the value of everything else in this question. The repository stops describing reality, recreating an environment produces something different from what was running, and the review process governs a fiction.

Detect it by comparing periodically and automatically, not by trusting that nobody clicked. Then decide per instance: either the change was wrong and reality should be corrected, or it was right and the code should be updated. Both are legitimate; leaving it is not.

Distinguish drift from expected variance. Autoscaled node counts and generated identifiers change without anyone editing anything, and a detector that reports them trains people to ignore the report.

On STACKIT. Drift detection is a property of your tooling rather than of the platform. A scheduled plan run in the pipeline that fails when it finds unexpected changes is the simplest version and usually sufficient.

The audit log is the complementary source: it records every action by users, service accounts and the platform, at organization, folder and project scope, and is enabled by default. When drift is found, that is where you look for what changed and who changed it. Note the 90-day retention in the Portal, which bounds how far back an investigation can reach unless the records are exported. See OPS 7.4.

Tradeoffs. Operational Excellence. A drift detector is a thing to operate and tune, and a noisy one is worse than none. Cost Optimization. Scheduled plan runs consume pipeline capacity.

Verify. When was drift last detected in your production environment, how was it found, and what was done about it? If the answer is that it has never been detected, is that because there is none?


OPS 3.4 Codify emergency changes afterwards

Section titled “OPS 3.4 Codify emergency changes afterwards”

Risk if not established: Medium

Sometimes the right thing is to fix production now. A rule that forbids it produces either a slower incident response or a rule everyone ignores, and the second is worse because it also removes the record.

Permit the emergency change and require the follow-up. The change is made, the incident ends, and then the same change is expressed in code so that the repository and reality agree again. Without that second step, the drift is permanent and it accumulates precisely in the components that have the most incidents.

Make the follow-up a tracked item with an owner and a deadline, in the same way incident actions are handled in OPS 9.4. An intention to codify it later has the same completion rate as any other unowned intention.

Watch the pattern rather than the individual instance. A component that repeatedly needs emergency changes is telling you something about its design or its automation, which is OPS 10.3.

On STACKIT. Emergency access through the STACKIT Portal or the CLI is recorded in the audit log by default, which gives the follow-up a factual basis: what was actually changed, by whom, and when. That is more reliable than reconstructing it from memory after a long night.

Where emergency access requires elevated permissions, the break-glass treatment in SEC 5 applies: tightly bounded, heavily logged, and reviewed after every use.

Tradeoffs. Security. An emergency path that bypasses review is a real exposure, which is why it is bounded and logged rather than removed. The tension is named in the Security tradeoffs.

Verify. For the last emergency change made outside the normal process, was it subsequently expressed in code, and how long did that take?


  • OPS 2 Development standards, which the pipeline also enforces
  • OPS 4 Deployment automation, which runs on the same definitions
  • OPS 6 Environment consistency, which is impossible without this
  • REL 9.2 Disaster recovery, whose sequence is executable only if the environment is codified
  • SEC 2 Segmentation and COST 2 Cost attribution, which share the resource hierarchy