Zum Inhalt springen
Beta

OPS 5. How do you limit the exposure of a bad change?

Zuletzt aktualisiert am

Testing reduces the probability that a change is bad. It does not reach zero, and the changes that get through are by definition the ones your tests did not anticipate.

Safe deployment accepts that and limits the consequence instead. A bad change reaches a small fraction of traffic, is detected by a signal rather than by a person, and is withdrawn automatically. The difference between that and a full rollout is the difference between a blip and an incident.

  • OPS 5.1 Expose changes progressively rather than all at once
  • OPS 5.2 Gate progression on health signals rather than on elapsed time
  • OPS 5.3 Roll back automatically when a gate fails
  • OPS 5.4 Advance one environment at a time, and one region at a time

OPS 5.1 Expose changes progressively rather than all at once

Section titled “OPS 5.1 Expose changes progressively rather than all at once”

Risk if not established: High

The mechanisms differ in cost and in what they let you observe:

Rolling update. Instances replaced in batches. The cheapest option and the usual default. It limits exposure over time rather than by audience, and it requires both versions to coexist briefly, which constrains schema and API compatibility exactly as OPS 4.2 describes.

Canary. A small share of traffic to the new version while the rest stays on the old. Gives a clean comparison between the two, which is what makes automated gating possible.

Blue-green. Two complete environments, traffic switched between them. Fast to reverse and expensive, because you run two of everything during the change.

Choose per flow, using the ranking from REL 2.2. A rolling update is proportionate for most things; a canary earns its complexity on the flows where a bad change is costly.

Whatever the mechanism, both versions run at once. That is a design constraint on the application and on the data, not an implementation detail, and it is the most common reason a progressive rollout has to be abandoned mid-flight.

On STACKIT. Kubernetes Engine gives you the standard Kubernetes rolling update behaviour by default, and canary or blue-green patterns are implemented with the routing tools you choose to run in the cluster. Cloud Foundry has its own application lifecycle for pushing new versions.

At the edge, load balancing is where traffic splitting between two backends is expressed if you are doing it outside the cluster.

Weighted distribution between target pools is what a platform-level canary needs at that position, so confirm it for the load balancer you are using before a rollout depends on it. In-cluster routing is the alternative that does not, which makes it the safer thing to design first.

Tradeoffs. Cost Optimization. Blue-green doubles the running environment during a deployment, and canaries mean operating two versions with two sets of telemetry. Operational Excellence. Every mechanism beyond a rolling update is a thing to configure, understand and debug under pressure.

Verify. For your most critical flow, what fraction of users sees a new version first, and for how long before it goes wider?


OPS 5.2 Gate progression on health signals rather than on elapsed time

Section titled “OPS 5.2 Gate progression on health signals rather than on elapsed time”

Risk if not established: High

Waiting ten minutes between stages is not a gate. It delays the rollout without deciding anything, and it will proceed just as happily through a version that is failing.

A real gate compares signals against a threshold and stops if they are not met. The signals that work are the same ones REL 10.2 alerts on, because they are the ones that reflect user experience: error rate, latency percentiles, and a small number of business-level indicators such as completed checkouts.

Compare the new version against the old rather than against an absolute threshold. Absolute thresholds fail in both directions: they trigger during unrelated load spikes and they miss regressions that stay inside a generous limit. A canary that is measurably worse than its control is a signal regardless of whether either has breached a limit.

Give the gate enough traffic and enough time to be meaningful. A canary receiving two requests a minute cannot distinguish a regression from noise, and a gate evaluated over thirty seconds will miss anything that appears under sustained load.

On STACKIT. Observability holds the metrics a gate queries, and its alerting is the mechanism for expressing the thresholds.

The query from a pipeline into those metrics is yours to build, since the gate is a step in STACKIT Pipelines rather than a platform feature. That is the usual division: the platform provides the signal store and the pipeline, and the policy that connects them is the part that encodes your judgement.

Which signals exist at all depends on OPS 7. A gate cannot query a metric that nobody emits.

Tradeoffs. Operational Excellence. Thresholds need tuning, and a gate that is too sensitive blocks good deployments while one that is too tolerant passes bad ones. Both need production data to correct, so expect to adjust rather than to get it right first.

Verify. What must be true for a deployment to advance to the next stage? Is that a measured condition or an elapsed time?


OPS 5.3 Roll back automatically when a gate fails

Section titled “OPS 5.3 Roll back automatically when a gate fails”

Risk if not established: High

A gate that pages a human has converted an automated safety mechanism into a manual one, at the time of day when humans are least available. The point of the gate is that the response does not wait for anyone.

Automatic rollback needs three things to be safe: a deterministic previous artefact from OPS 4.1, a rollback path exercised often enough to be trusted from OPS 4.2, and backwards-compatible data changes so that reverting the code does not strand the schema.

Bound the automation. A rollback loop that redeploys and reverts repeatedly is worse than stopping, so a single automatic rollback followed by a halt and an alert is the right shape. Automation failing safely means stopping, not retrying.

Notify afterwards rather than asking permission beforehand. The record of what happened, which gate failed and what the signal looked like is what the investigation needs, and it belongs in the same place as incident records under OPS 9.

On STACKIT. The rollback step lives in your pipeline and its mechanics depend on the runtime, as in OPS 4.2. Kubernetes revision rollback and Cloud Foundry’s application lifecycle both support it natively; a Compute Engine deployment reverts by redeploying the previous artefact.

Automatic rollback needs the pipeline to hold deployment credentials, which is OPS 4.1 again and another reason the pipeline identity belongs under SEC 5.

Tradeoffs. Reliability. An automated action with a bad condition applies its mistake at machine speed, which is why the bound matters. A rollback triggered by a monitoring failure rather than by an application failure is the specific case to guard against.

Verify. What happens when a health gate fails: automatic rollback, a page, or nothing? If automatic, when did it last trigger and was the outcome correct?


OPS 5.4 Advance one environment at a time, and one region at a time

Section titled “OPS 5.4 Advance one environment at a time, and one region at a time”

Risk if not established: High

Deploying everywhere simultaneously means a bad change reaches everything simultaneously, which removes the containment that having separate environments was supposed to provide.

Advance in order, with the earlier stages carrying real signal. That requires environments that are similar enough for the earlier result to predict the later one, which is OPS 6, and it requires that a failure in an earlier stage actually stops the promotion rather than being waved through.

Where a workload runs in more than one region, treat them as sequential stages for the same reason. A change that passes in the first region and fails in the second has told you something valuable, and it has told you at half the blast radius.

Give each stage a bake time proportional to what it can detect. Some failure modes appear only under sustained production load or at a daily peak, and promoting through all stages in an hour means none of them were exercised.

On STACKIT. Separation between environments is expressed through the Resource Manager hierarchy: distinct projects per environment give distinct access, distinct quotas and distinct billing, which is the same structure SEC 2 and COST 2 ask for and the reason OPS 6 becomes practical.

Regional sequencing across eu01 and eu02 follows from regions and availability zones . If a workload is single-region, this reduces to environment sequencing.

Tradeoffs. Operational Excellence. More stages means a longer path from merge to production, which pushes against OPS 4.3. Resolve by making each stage fast rather than by removing stages, and by scoping the number of stages to what each one genuinely detects.

Verify. How long does a change take to travel from merge to full production, and how many independent stages does it pass? Has any stage ever stopped a promotion?


  • OPS 4 Deployment automation, which supplies the artefact and the rollback path
  • OPS 6 Environment consistency, without which earlier stages predict nothing
  • OPS 7 Observability, which supplies the signals the gates query
  • REL 10.2 Alerting on symptoms, which uses the same signals
  • REL 2.2 Flow ranking, which decides where the expensive mechanisms are justified