---
id: OPS05
pillar: operational-excellence
title: OPS 5. How do you limit the exposure of a bad change?
description: A deployment strategy that depends on someone watching a dashboard fails at three in the morning. Progressive exposure, health gates and automatic rollback.
status: draft
services: [kubernetes-engine, cloud-foundry, observability]
sidebar:
  order: 14
  label: Safe deployment
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/ops-05-safe-deployment/"
source_file: "docs/architecture/pillars/operational-excellence/ops-05-safe-deployment.mdx"
---

Testing reduces the probability that a change is bad. It does not reach zero, and the changes that
get through are by definition the ones your tests did not anticipate.

Safe deployment accepts that and limits the consequence instead. A bad change reaches a small
fraction of traffic, is detected by a signal rather than by a person, and is withdrawn
automatically. The difference between that and a full rollout is the difference between a blip and
an incident.

## Best practices

- [`OPS 5.1`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-51-expose-changes-progressively-rather-than-all-at-once) Expose changes progressively rather than all at once
- [`OPS 5.2`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-52-gate-progression-on-health-signals-rather-than-on-elapsed-time) Gate progression on health signals rather than on elapsed time
- [`OPS 5.3`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-53-roll-back-automatically-when-a-gate-fails) Roll back automatically when a gate fails
- [`OPS 5.4`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-54-advance-one-environment-at-a-time-and-one-region-at-a-time) Advance one environment at a time, and one region at a time

---

## OPS 5.1 Expose changes progressively rather than all at once

**Risk if not established:** High

The mechanisms differ in cost and in what they let you observe:

**Rolling update.** Instances replaced in batches. The cheapest option and the usual default. It
limits exposure over time rather than by audience, and it requires both versions to coexist
briefly, which constrains schema and API compatibility exactly as [`OPS 4.2`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-42-make-rollback-an-ordinary-operation-and-exercise-it) describes.

**Canary.** A small share of traffic to the new version while the rest stays on the old. Gives a
clean comparison between the two, which is what makes automated gating possible.

**Blue-green.** Two complete environments, traffic switched between them. Fast to reverse and
expensive, because you run two of everything during the change.

Choose per flow, using the ranking from [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume). A rolling update is proportionate for most
things; a canary earns its complexity on the flows where a bad change is costly.

Whatever the mechanism, both versions run at once. That is a design constraint on the application
and on the data, not an implementation detail, and it is the most common reason a progressive
rollout has to be abandoned mid-flight.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/">Kubernetes
Engine</LinkChip> gives you the standard
Kubernetes rolling update behaviour by default, and canary or blue-green patterns are implemented
with the routing tools you choose to run in the cluster. <LinkChip href="https://docs.stackit.cloud/products/runtime/cloud-foundry/">Cloud
Foundry</LinkChip> has its own application
lifecycle for pushing new versions.

At the edge,
<LinkChip href="https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/">load balancing</LinkChip>
is where traffic splitting between two backends is expressed if you are doing it outside the
cluster.

Weighted distribution between target pools is what a platform-level canary needs at that
position, so confirm it for the load balancer you are using before a rollout depends on it.
In-cluster routing is the alternative that does not, which makes it the safer thing to design
first.

**Tradeoffs.** **Cost Optimization.** Blue-green doubles the running environment during a
deployment, and canaries mean operating two versions with two sets of telemetry.
**Operational Excellence.** Every mechanism beyond a rolling update is a thing to configure,
understand and debug under pressure.

**Verify.** For your most critical flow, what fraction of users sees a new version first, and for
how long before it goes wider?

---

## OPS 5.2 Gate progression on health signals rather than on elapsed time

**Risk if not established:** High

Waiting ten minutes between stages is not a gate. It delays the rollout without deciding anything,
and it will proceed just as happily through a version that is failing.

A real gate compares signals against a threshold and stops if they are not met. The signals that
work are the same ones [`REL 10.2`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-102-alert-on-user-visible-symptoms-rather-than-on-component-metrics-alone) alerts on, because they are the ones that reflect user
experience: error rate, latency percentiles, and a small number of business-level indicators such
as completed checkouts.

Compare the new version against the old rather than against an absolute threshold. Absolute
thresholds fail in both directions: they trigger during unrelated load spikes and they miss
regressions that stay inside a generous limit. A canary that is measurably worse than its control
is a signal regardless of whether either has breached a limit.

Give the gate enough traffic and enough time to be meaningful. A canary receiving two requests a
minute cannot distinguish a regression from noise, and a gate evaluated over thirty seconds will
miss anything that appears under sustained load.

**On STACKIT.**
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> holds
the metrics a gate queries, and its
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/getting-started/alerting-overview/">alerting</LinkChip>
is the mechanism for expressing the thresholds.

The query from a pipeline into those metrics is yours to build, since the gate is a step in
<LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/basics/stackit-pipelines/">STACKIT
Pipelines</LinkChip>
rather than a platform feature. That is the usual division: the platform provides the signal store
and the pipeline, and the policy that connects them is the part that encodes your judgement.

Which signals exist at all depends on [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/). A gate cannot query a metric that nobody emits.

**Tradeoffs.** **Operational Excellence.** Thresholds need tuning, and a gate that is too
sensitive blocks good deployments while one that is too tolerant passes bad ones. Both need
production data to correct, so expect to adjust rather than to get it right first.

**Verify.** What must be true for a deployment to advance to the next stage? Is that a measured
condition or an elapsed time?

---

## OPS 5.3 Roll back automatically when a gate fails

**Risk if not established:** High

A gate that pages a human has converted an automated safety mechanism into a manual one, at the
time of day when humans are least available. The point of the gate is that the response does not
wait for anyone.

Automatic rollback needs three things to be safe: a deterministic previous artefact from `OPS
4.1`, a rollback path exercised often enough to be trusted from [`OPS 4.2`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-42-make-rollback-an-ordinary-operation-and-exercise-it), and
backwards-compatible data changes so that reverting the code does not strand the schema.

Bound the automation. A rollback loop that redeploys and reverts repeatedly is worse than
stopping, so a single automatic rollback followed by a halt and an alert is the right shape.
Automation failing safely means stopping, not retrying.

Notify afterwards rather than asking permission beforehand. The record of what happened, which
gate failed and what the signal looked like is what the investigation needs, and it belongs in the
same place as incident records under [`OPS 9`](/architecture/pillars/operational-excellence/ops-09-incident-management/).

**On STACKIT.** The rollback step lives in your pipeline and its mechanics depend on the runtime,
as in [`OPS 4.2`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-42-make-rollback-an-ordinary-operation-and-exercise-it). Kubernetes revision rollback and Cloud Foundry's application lifecycle both
support it natively; a Compute Engine deployment reverts by redeploying the previous artefact.

Automatic rollback needs the pipeline to hold deployment credentials, which is [`OPS 4.1`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-41-build-one-automated-path-from-source-to-production-used-by-everyone) again and
another reason the pipeline identity belongs under [`SEC 5`](/architecture/pillars/security/sec-05-least-privilege/).

**Tradeoffs.** **Reliability.** An automated action with a bad condition applies its mistake at
machine speed, which is why the bound matters. A rollback triggered by a monitoring failure rather
than by an application failure is the specific case to guard against.

**Verify.** What happens when a health gate fails: automatic rollback, a page, or nothing? If
automatic, when did it last trigger and was the outcome correct?

---

## OPS 5.4 Advance one environment at a time, and one region at a time

**Risk if not established:** High

Deploying everywhere simultaneously means a bad change reaches everything simultaneously, which
removes the containment that having separate environments was supposed to provide.

Advance in order, with the earlier stages carrying real signal. That requires environments that
are similar enough for the earlier result to predict the later one, which is [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/), and it
requires that a failure in an earlier stage actually stops the promotion rather than being waved
through.

Where a workload runs in more than one region, treat them as sequential stages for the same
reason. A change that passes in the first region and fails in the second has told you something
valuable, and it has told you at half the blast radius.

Give each stage a bake time proportional to what it can detect. Some failure modes appear only
under sustained production load or at a daily peak, and promoting through all stages in an hour
means none of them were exercised.

**On STACKIT.** Separation between environments is expressed through the <LinkChip href="https://docs.stackit.cloud/platform/resource-manager/">Resource
Manager</LinkChip>
hierarchy: distinct projects per environment give
distinct access, distinct quotas and distinct billing, which is the same structure [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) and
[`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) ask for and the reason [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) becomes practical.

Regional sequencing across `eu01` and `eu02` follows from
<LinkChip href="https://docs.stackit.cloud/platform/regions/">regions and availability zones</LinkChip>. If a workload is
single-region, this reduces to environment sequencing.

**Tradeoffs.** **Operational Excellence.** More stages means a longer path from merge to
production, which pushes against [`OPS 4.3`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-43-keep-changes-small-and-deploy-frequently). Resolve by making each stage fast rather than by
removing stages, and by scoping the number of stages to what each one genuinely detects.

**Verify.** How long does a change take to travel from merge to full production, and how many
independent stages does it pass? Has any stage ever stopped a promotion?

---

## Related

- [`OPS 4`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/) Deployment automation, which supplies the artefact and the rollback path
- [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) Environment consistency, without which earlier stages predict nothing
- [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) Observability, which supplies the signals the gates query
- [`REL 10.2`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-102-alert-on-user-visible-symptoms-rather-than-on-component-metrics-alone) Alerting on symptoms, which uses the same signals
- [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume) Flow ranking, which decides where the expensive mechanisms are justified
