OPS 5. How do you limit the exposure of a bad change?
Last updated on
Testing reduces the probability that a change is bad. It does not reach zero, and the changes that get through are by definition the ones your tests did not anticipate.
Safe deployment accepts that and limits the consequence instead. A bad change reaches a small fraction of traffic, is detected by a signal rather than by a person, and is withdrawn automatically. The difference between that and a full rollout is the difference between a blip and an incident.
Best practices
Section titled “Best practices”OPS 5.1Expose changes progressively rather than all at onceOPS 5.2Gate progression on health signals rather than on elapsed timeOPS 5.3Roll back automatically when a gate failsOPS 5.4Advance one environment at a time, and one region at a time
OPS 5.1 Expose changes progressively rather than all at once
Section titled “OPS 5.1 Expose changes progressively rather than all at once”Risk if not established: High
The mechanisms differ in cost and in what they let you observe:
Rolling update. Instances replaced in batches. The cheapest option and the usual default. It
limits exposure over time rather than by audience, and it requires both versions to coexist
briefly, which constrains schema and API compatibility exactly as OPS 4.2 describes.
Canary. A small share of traffic to the new version while the rest stays on the old. Gives a clean comparison between the two, which is what makes automated gating possible.
Blue-green. Two complete environments, traffic switched between them. Fast to reverse and expensive, because you run two of everything during the change.
Choose per flow, using the ranking from REL 2.2. A rolling update is proportionate for most
things; a canary earns its complexity on the flows where a bad change is costly.
Whatever the mechanism, both versions run at once. That is a design constraint on the application and on the data, not an implementation detail, and it is the most common reason a progressive rollout has to be abandoned mid-flight.
On STACKIT. Kubernetes Engine gives you the standard Kubernetes rolling update behaviour by default, and canary or blue-green patterns are implemented with the routing tools you choose to run in the cluster. Cloud Foundry has its own application lifecycle for pushing new versions.
At the edge, load balancing is where traffic splitting between two backends is expressed if you are doing it outside the cluster.
Weighted distribution between target pools is what a platform-level canary needs at that position, so confirm it for the load balancer you are using before a rollout depends on it. In-cluster routing is the alternative that does not, which makes it the safer thing to design first.
Tradeoffs. Cost Optimization. Blue-green doubles the running environment during a deployment, and canaries mean operating two versions with two sets of telemetry. Operational Excellence. Every mechanism beyond a rolling update is a thing to configure, understand and debug under pressure.
Verify. For your most critical flow, what fraction of users sees a new version first, and for how long before it goes wider?
OPS 5.2 Gate progression on health signals rather than on elapsed time
Section titled “OPS 5.2 Gate progression on health signals rather than on elapsed time”Risk if not established: High
Waiting ten minutes between stages is not a gate. It delays the rollout without deciding anything, and it will proceed just as happily through a version that is failing.
A real gate compares signals against a threshold and stops if they are not met. The signals that
work are the same ones REL 10.2 alerts on, because they are the ones that reflect user
experience: error rate, latency percentiles, and a small number of business-level indicators such
as completed checkouts.
Compare the new version against the old rather than against an absolute threshold. Absolute thresholds fail in both directions: they trigger during unrelated load spikes and they miss regressions that stay inside a generous limit. A canary that is measurably worse than its control is a signal regardless of whether either has breached a limit.
Give the gate enough traffic and enough time to be meaningful. A canary receiving two requests a minute cannot distinguish a regression from noise, and a gate evaluated over thirty seconds will miss anything that appears under sustained load.
On STACKIT. Observability holds the metrics a gate queries, and its alerting is the mechanism for expressing the thresholds.
The query from a pipeline into those metrics is yours to build, since the gate is a step in STACKIT Pipelines rather than a platform feature. That is the usual division: the platform provides the signal store and the pipeline, and the policy that connects them is the part that encodes your judgement.
Which signals exist at all depends on OPS 7. A gate cannot query a metric that nobody emits.
Tradeoffs. Operational Excellence. Thresholds need tuning, and a gate that is too sensitive blocks good deployments while one that is too tolerant passes bad ones. Both need production data to correct, so expect to adjust rather than to get it right first.
Verify. What must be true for a deployment to advance to the next stage? Is that a measured condition or an elapsed time?
OPS 5.3 Roll back automatically when a gate fails
Section titled “OPS 5.3 Roll back automatically when a gate fails”Risk if not established: High
A gate that pages a human has converted an automated safety mechanism into a manual one, at the time of day when humans are least available. The point of the gate is that the response does not wait for anyone.
Automatic rollback needs three things to be safe: a deterministic previous artefact from OPS 4.1, a rollback path exercised often enough to be trusted from OPS 4.2, and
backwards-compatible data changes so that reverting the code does not strand the schema.
Bound the automation. A rollback loop that redeploys and reverts repeatedly is worse than stopping, so a single automatic rollback followed by a halt and an alert is the right shape. Automation failing safely means stopping, not retrying.
Notify afterwards rather than asking permission beforehand. The record of what happened, which
gate failed and what the signal looked like is what the investigation needs, and it belongs in the
same place as incident records under OPS 9.
On STACKIT. The rollback step lives in your pipeline and its mechanics depend on the runtime,
as in OPS 4.2. Kubernetes revision rollback and Cloud Foundry’s application lifecycle both
support it natively; a Compute Engine deployment reverts by redeploying the previous artefact.
Automatic rollback needs the pipeline to hold deployment credentials, which is OPS 4.1 again and
another reason the pipeline identity belongs under SEC 5.
Tradeoffs. Reliability. An automated action with a bad condition applies its mistake at machine speed, which is why the bound matters. A rollback triggered by a monitoring failure rather than by an application failure is the specific case to guard against.
Verify. What happens when a health gate fails: automatic rollback, a page, or nothing? If automatic, when did it last trigger and was the outcome correct?
OPS 5.4 Advance one environment at a time, and one region at a time
Section titled “OPS 5.4 Advance one environment at a time, and one region at a time”Risk if not established: High
Deploying everywhere simultaneously means a bad change reaches everything simultaneously, which removes the containment that having separate environments was supposed to provide.
Advance in order, with the earlier stages carrying real signal. That requires environments that
are similar enough for the earlier result to predict the later one, which is OPS 6, and it
requires that a failure in an earlier stage actually stops the promotion rather than being waved
through.
Where a workload runs in more than one region, treat them as sequential stages for the same reason. A change that passes in the first region and fails in the second has told you something valuable, and it has told you at half the blast radius.
Give each stage a bake time proportional to what it can detect. Some failure modes appear only under sustained production load or at a daily peak, and promoting through all stages in an hour means none of them were exercised.
On STACKIT. Separation between environments is expressed through the Resource
Manager
hierarchy: distinct projects per environment give
distinct access, distinct quotas and distinct billing, which is the same structure SEC 2 and
COST 2 ask for and the reason OPS 6 becomes practical.
Regional sequencing across eu01 and eu02 follows from
regions and availability zones . If a workload is
single-region, this reduces to environment sequencing.
Tradeoffs. Operational Excellence. More stages means a longer path from merge to
production, which pushes against OPS 4.3. Resolve by making each stage fast rather than by
removing stages, and by scoping the number of stages to what each one genuinely detects.
Verify. How long does a change take to travel from merge to full production, and how many independent stages does it pass? Has any stage ever stopped a promotion?
Related
Section titled “Related”OPS 4Deployment automation, which supplies the artefact and the rollback pathOPS 6Environment consistency, without which earlier stages predict nothingOPS 7Observability, which supplies the signals the gates queryREL 10.2Alerting on symptoms, which uses the same signalsREL 2.2Flow ranking, which decides where the expensive mechanisms are justified