Zum Inhalt springen
Beta

OPS 4. How do you make deployment repeatable and reversible?

Zuletzt aktualisiert am

The instinct that deployment is risky, and should therefore be rare and heavily ceremonied, is self-reinforcing and backwards. Rarity accumulates changes, large changes are harder to review and much harder to diagnose, and ceremony makes deployment expensive, which makes it rarer still. The risk was created by the caution.

The inversion is small, frequent, automated and reversible. Each deployment carries one comprehensible change, a failure has an obvious suspect, and rollback is routine rather than a decision requiring a meeting.

  • OPS 4.1 Build one automated path from source to production, used by everyone
  • OPS 4.2 Make rollback an ordinary operation and exercise it
  • OPS 4.3 Keep changes small and deploy frequently
  • OPS 4.4 Separate deploying from releasing

OPS 4.1 Build one automated path from source to production, used by everyone

Section titled “OPS 4.1 Build one automated path from source to production, used by everyone”

Risk if not established: High

One path, no exceptions. The moment there is a second way to get code into production, the guarantees of the first are gone: you no longer know what is running, the audit trail has holes, and the emergency path is the one least tested.

The urgent fix is the case that tests the principle. If the pipeline is too slow to use during an incident, the answer is to make it faster rather than to build a bypass. A pipeline that is bypassed under pressure is a pipeline that is absent when it matters most.

The path should be deterministic: the same commit produces the same artefact, and the artefact deployed to production is the one that was tested rather than a rebuild. Rebuilding per environment reintroduces the variance the pipeline existed to remove.

Deployment credentials belong to the pipeline rather than to people. That is what makes the single path enforceable, and it makes the pipeline identity a privileged one that SEC 4 and SEC 5 should be applied to.

On STACKIT. STACKIT Pipelines is the native option, built into STACKIT Git and compatible with GitHub Actions workflows, which means existing workflow definitions and community actions can be reused. Jobs run on managed or custom runners , and the first steps guide covers the basic shape.

Artefacts belong in Container Registry , which also scans them, so the same store serves deployment and SEC 10.

The deployment target shapes the mechanism. Cloud Foundry has an application push model, Kubernetes Engine uses whatever Kubernetes deployment approach you choose, and Compute Engine deployments are yours to orchestrate through the API, CLI or Terraform. All three are legitimate; the second and third require you to build more of the path.

Note that both the registry and the source live on the platform, which puts them on the recovery path analysed in REL 3.3 and REL 9.4.

Tradeoffs. Cost Optimization. Building the path is real up-front work, and pipeline capacity is a running cost. Security. A pipeline that can change production is a high-value target and needs production-grade access control, which is stated in the Operational Excellence tradeoffs.

Verify. How many ways are there to change what runs in production? For each, who can use it and what does it record?


OPS 4.2 Make rollback an ordinary operation and exercise it

Section titled “OPS 4.2 Make rollback an ordinary operation and exercise it”

Risk if not established: High

The ability to undo a deployment is what makes deploying safe. Without it, every release is a one-way door and the caution that follows is rational.

Rollback needs to be as automated as deployment, achievable without a decision meeting, and exercised often enough that people trust it. A rollback procedure first executed during an incident is a hypothesis.

The part that is genuinely hard is state. Application code rolls back cleanly; database schemas usually do not. The discipline is to make schema changes backwards compatible so that the previous application version still works against the new schema: add before removing, deploy in two steps, and never combine a destructive migration with a code change in one release.

Decide in advance what triggers a rollback rather than debating it while a system is degraded. An error rate threshold agreed beforehand converts a judgement call into an action, which is what OPS 5.3 automates.

On STACKIT. Rollback mechanics depend on the runtime. Kubernetes deployments support revision rollback natively; Cloud Foundry has its own application lifecycle. For Compute Engine, rollback usually means redeploying the previous artefact, which is why the deterministic artefact from OPS 4.1 matters.

For data, restoring is not rollback. A restore under backup and clone loses everything written since the backup, which is REL 8 rather than a deployment operation. Treating a restore as a rollback path is how a bad deployment becomes a data loss incident.

Tradeoffs. Performance Efficiency. Backwards-compatible schema changes mean two deployments where one would have done, and a period where the schema carries both shapes. That is the cost of being able to undo.

Verify. When did you last roll back a production deployment, how long did it take, and was a meeting required? What would happen if the change included a schema migration?


OPS 4.3 Keep changes small and deploy frequently

Section titled “OPS 4.3 Keep changes small and deploy frequently”

Risk if not established: Medium

The size of a deployment is the strongest predictor of how hard it is to diagnose when something breaks. One change has one suspect. Forty changes have forty, and the interaction between them.

Frequency and size are the same lever. Deploying more often necessarily means each deployment carries less, which is why teams that deploy daily recover faster than teams that deploy quarterly. The practice is exercised constantly rather than annually.

What blocks frequency is usually not technical. Manual approval steps that take days, test suites that take hours, and coordination requirements between teams all push toward batching. Each of those is addressable, and each is worth more than it looks because the benefit is per deployment.

Frequent deployment is only safe with the machinery of OPS 5. Without progressive exposure and automatic rollback, more deployments is simply more chances to break production. The order matters: build the safety, then increase the frequency.

On STACKIT. Nothing platform-specific determines your deployment frequency. What the platform affects is the friction: pipeline duration on managed or custom runners , and how long the deployment mechanism itself takes on your chosen runtime.

Where an approval step is genuinely required, for example under a regulatory obligation, manual workflow approval keeps it inside the automated path rather than outside it, which preserves the single path from OPS 4.1.

Aus der STACKIT-DokuPipelines – Manuelle Workflow-Genehmigung › EinschränkungenStand der Quelle 14.07.2026 · übernommen 05.10.2026
  • Während der Workflow pausiert ist, verbraucht er weiterhin eine Zuweisung für gleichzeitige Jobs aus dem maximalen Kontingent von 20 gleichzeitigen Jobs.
  • Ein pausierter Job führt weiterhin Berechnungen auf der Instanz oder der virtuellen Maschine aus, sodass fortlaufend Kosten anfallen.
  • Ablauf (wird auch an anderer Stelle in diesem Dokument erwähnt):
    • Ein Job (einschließlich eines pausierten Jobs) schlägt nach 4 Stunden fehl, und ein Workflow schlägt nach 35 Tagen fehl.
    • Token der StackitGit App laufen nach 1 Stunde ab. Das bedeutet, dass die Dauer für die Genehmigung 60 Minuten nicht überschreiten darf, da der Job andernfalls aufgrund ungültiger Zugangsdaten fehlschlägt.
Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

Tradeoffs. Reliability, in the short term. Change is the leading cause of incidents, and more deployments means more changes. The resolution is that the risk comes from unmanaged exposure rather than from frequency, which is exactly what OPS 5 addresses.

Verify. How often do you deploy to production, and how many changes does a typical deployment contain? What is the longest step in the path from merge to production?


Risk if not established: Medium

Deploying puts code in production. Releasing makes it visible to users. Conflating them means every user-facing change requires a deployment, and every deployment carries user-facing risk.

Separating them with feature flags decouples the two: code ships dark, is enabled for a small group, then widened, and can be disabled without a deployment. That last property is what makes the response to a bad feature seconds rather than a deployment cycle.

The cost is real and worth stating. Flags accumulate, every one is a branch in the code, and a codebase with hundreds of stale flags is harder to reason about than one without any. Flags need owners and removal dates, and removing them is work that competes with features.

Use it where the risk justifies the complexity: user-facing changes on ranked flows, migrations that need a gradual switch, anything you might want to disable without waiting for a pipeline. Not for every change.

On STACKIT. Feature flag management is an application concern, and there is no managed feature flag service. The usual options are a library with configuration, a self-operated open-source service, or a third party. The third option is worth checking against SOV 6, since a flag service sits in the request path and sees traffic.

Where flags are held in configuration rather than in a dedicated service, Secrets Manager is for secrets rather than for feature configuration, so the two should not be conflated.

Tradeoffs. Operational Excellence. Flags are permanent complexity unless actively removed, and the removal is the part that gets skipped. Performance Efficiency. A flag evaluated on a hot path costs a lookup, which is small until it is not.

Verify. Can you disable a user-facing feature without deploying? How many feature flags exist in your codebase, and how many have an owner and a removal date?


  • OPS 3 Everything as code, which the pipeline deploys from
  • OPS 5 Safe deployment, which is what makes frequent deployment safe
  • OPS 6 Environment consistency, so that testing predicts production
  • REL 8 Backup and restore, which is not a rollback mechanism
  • SEC 10 Supply chain, which the artefact path shares