OPS 4. How do you make deployment repeatable and reversible?
Zuletzt aktualisiert am
The instinct that deployment is risky, and should therefore be rare and heavily ceremonied, is self-reinforcing and backwards. Rarity accumulates changes, large changes are harder to review and much harder to diagnose, and ceremony makes deployment expensive, which makes it rarer still. The risk was created by the caution.
The inversion is small, frequent, automated and reversible. Each deployment carries one comprehensible change, a failure has an obvious suspect, and rollback is routine rather than a decision requiring a meeting.
Best practices
Section titled “Best practices”OPS 4.1Build one automated path from source to production, used by everyoneOPS 4.2Make rollback an ordinary operation and exercise itOPS 4.3Keep changes small and deploy frequentlyOPS 4.4Separate deploying from releasing
OPS 4.1 Build one automated path from source to production, used by everyone
Section titled “OPS 4.1 Build one automated path from source to production, used by everyone”Risk if not established: High
One path, no exceptions. The moment there is a second way to get code into production, the guarantees of the first are gone: you no longer know what is running, the audit trail has holes, and the emergency path is the one least tested.
The urgent fix is the case that tests the principle. If the pipeline is too slow to use during an incident, the answer is to make it faster rather than to build a bypass. A pipeline that is bypassed under pressure is a pipeline that is absent when it matters most.
The path should be deterministic: the same commit produces the same artefact, and the artefact deployed to production is the one that was tested rather than a rebuild. Rebuilding per environment reintroduces the variance the pipeline existed to remove.
Deployment credentials belong to the pipeline rather than to people. That is what makes the single
path enforceable, and it makes the pipeline identity a privileged one that SEC 4 and SEC 5
should be applied to.
On STACKIT. STACKIT Pipelines is the native option, built into STACKIT Git and compatible with GitHub Actions workflows, which means existing workflow definitions and community actions can be reused. Jobs run on managed or custom runners , and the first steps guide covers the basic shape.
Artefacts belong in
Container Registry ,
which also scans them, so the same store serves deployment and SEC 10.
The deployment target shapes the mechanism. Cloud Foundry has an application push model, Kubernetes Engine uses whatever Kubernetes deployment approach you choose, and Compute Engine deployments are yours to orchestrate through the API, CLI or Terraform. All three are legitimate; the second and third require you to build more of the path.
Note that both the registry and the source live on the platform, which puts them on the recovery
path analysed in REL 3.3 and REL 9.4.
Tradeoffs. Cost Optimization. Building the path is real up-front work, and pipeline capacity is a running cost. Security. A pipeline that can change production is a high-value target and needs production-grade access control, which is stated in the Operational Excellence tradeoffs.
Verify. How many ways are there to change what runs in production? For each, who can use it and what does it record?
OPS 4.2 Make rollback an ordinary operation and exercise it
Section titled “OPS 4.2 Make rollback an ordinary operation and exercise it”Risk if not established: High
The ability to undo a deployment is what makes deploying safe. Without it, every release is a one-way door and the caution that follows is rational.
Rollback needs to be as automated as deployment, achievable without a decision meeting, and exercised often enough that people trust it. A rollback procedure first executed during an incident is a hypothesis.
The part that is genuinely hard is state. Application code rolls back cleanly; database schemas usually do not. The discipline is to make schema changes backwards compatible so that the previous application version still works against the new schema: add before removing, deploy in two steps, and never combine a destructive migration with a code change in one release.
Decide in advance what triggers a rollback rather than debating it while a system is degraded. An
error rate threshold agreed beforehand converts a judgement call into an action, which is what
OPS 5.3 automates.
On STACKIT. Rollback mechanics depend on the runtime. Kubernetes deployments support revision
rollback natively; Cloud Foundry has
its own application lifecycle. For Compute Engine, rollback usually means redeploying the previous
artefact, which is why the deterministic artefact from OPS 4.1 matters.
For data, restoring is not rollback. A restore under backup and
clone
loses everything written since the backup, which is REL 8 rather than a deployment operation.
Treating a restore as a rollback path is how a bad deployment becomes a data loss incident.
Tradeoffs. Performance Efficiency. Backwards-compatible schema changes mean two deployments where one would have done, and a period where the schema carries both shapes. That is the cost of being able to undo.
Verify. When did you last roll back a production deployment, how long did it take, and was a meeting required? What would happen if the change included a schema migration?
OPS 4.3 Keep changes small and deploy frequently
Section titled “OPS 4.3 Keep changes small and deploy frequently”Risk if not established: Medium
The size of a deployment is the strongest predictor of how hard it is to diagnose when something breaks. One change has one suspect. Forty changes have forty, and the interaction between them.
Frequency and size are the same lever. Deploying more often necessarily means each deployment carries less, which is why teams that deploy daily recover faster than teams that deploy quarterly. The practice is exercised constantly rather than annually.
What blocks frequency is usually not technical. Manual approval steps that take days, test suites that take hours, and coordination requirements between teams all push toward batching. Each of those is addressable, and each is worth more than it looks because the benefit is per deployment.
Frequent deployment is only safe with the machinery of OPS 5. Without progressive exposure and
automatic rollback, more deployments is simply more chances to break production. The order
matters: build the safety, then increase the frequency.
On STACKIT. Nothing platform-specific determines your deployment frequency. What the platform affects is the friction: pipeline duration on managed or custom runners , and how long the deployment mechanism itself takes on your chosen runtime.
Where an approval step is genuinely required, for example under a regulatory obligation, manual
workflow
approval
keeps it inside the automated path rather than outside it, which preserves the single path from
OPS 4.1.
- Während der Workflow pausiert ist, verbraucht er weiterhin eine Zuweisung für gleichzeitige Jobs aus dem maximalen Kontingent von 20 gleichzeitigen Jobs.
- Ein pausierter Job führt weiterhin Berechnungen auf der Instanz oder der virtuellen Maschine aus, sodass fortlaufend Kosten anfallen.
- Ablauf (wird auch an anderer Stelle in diesem Dokument erwähnt):
- Ein Job (einschließlich eines pausierten Jobs) schlägt nach 4 Stunden fehl, und ein Workflow schlägt nach 35 Tagen fehl.
- Token der StackitGit App laufen nach 1 Stunde ab. Das bedeutet, dass die Dauer für die Genehmigung 60 Minuten nicht überschreiten darf, da der Job andernfalls aufgrund ungültiger Zugangsdaten fehlschlägt.
Was ist das?
Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.
Tradeoffs. Reliability, in the short term. Change is the leading cause of incidents, and
more deployments means more changes. The resolution is that the risk comes from unmanaged exposure
rather than from frequency, which is exactly what OPS 5 addresses.
Verify. How often do you deploy to production, and how many changes does a typical deployment contain? What is the longest step in the path from merge to production?
OPS 4.4 Separate deploying from releasing
Section titled “OPS 4.4 Separate deploying from releasing”Risk if not established: Medium
Deploying puts code in production. Releasing makes it visible to users. Conflating them means every user-facing change requires a deployment, and every deployment carries user-facing risk.
Separating them with feature flags decouples the two: code ships dark, is enabled for a small group, then widened, and can be disabled without a deployment. That last property is what makes the response to a bad feature seconds rather than a deployment cycle.
The cost is real and worth stating. Flags accumulate, every one is a branch in the code, and a codebase with hundreds of stale flags is harder to reason about than one without any. Flags need owners and removal dates, and removing them is work that competes with features.
Use it where the risk justifies the complexity: user-facing changes on ranked flows, migrations that need a gradual switch, anything you might want to disable without waiting for a pipeline. Not for every change.
On STACKIT. Feature flag management is an application concern, and there is no managed feature
flag service. The usual options are a library with configuration, a self-operated
open-source service, or a third party. The third option is worth checking against SOV 6, since a
flag service sits in the request path and sees traffic.
Where flags are held in configuration rather than in a dedicated service, Secrets Manager is for secrets rather than for feature configuration, so the two should not be conflated.
Tradeoffs. Operational Excellence. Flags are permanent complexity unless actively removed, and the removal is the part that gets skipped. Performance Efficiency. A flag evaluated on a hot path costs a lookup, which is small until it is not.
Verify. Can you disable a user-facing feature without deploying? How many feature flags exist in your codebase, and how many have an owner and a removal date?
Related
Section titled “Related”OPS 3Everything as code, which the pipeline deploys fromOPS 5Safe deployment, which is what makes frequent deployment safeOPS 6Environment consistency, so that testing predicts productionREL 8Backup and restore, which is not a rollback mechanismSEC 10Supply chain, which the artefact path shares