Zum Inhalt springen
Beta

OPS 6. How do you keep environments consistent from development to production?

Zuletzt aktualisiert am

Every other question in this pillar assumes that what happens in a pre-production environment predicts what will happen in production. When environments diverge, that assumption quietly stops holding, and the tests, the deployment stages and the rehearsals all keep running while telling you less and less.

Divergence is never decided. It accumulates: a version upgraded in one place, a component simplified because the full one was expensive, a configuration mechanism that differs because someone was in a hurry.

  • OPS 6.1 Let environments differ in scale and data, not in shape
  • OPS 6.2 Build every environment from the same definitions
  • OPS 6.3 Keep platform versions aligned, and upgrade in the same order as you deploy
  • OPS 6.4 Test with production-like data without using production data

OPS 6.1 Let environments differ in scale and data, not in shape

Section titled “OPS 6.1 Let environments differ in scale and data, not in shape”

Risk if not established: High

Shape is topology, configuration mechanism, component set and versions. Scale is instance sizes, replica counts and data volume. Differences in scale are expected and manageable. Differences in shape are what make a pre-production result meaningless.

The substitutions that cause the most trouble are the ones that feel harmless. Running a single database instance instead of a replica set means the failover behaviour is untested. Using an in-memory queue instead of the real broker means the delivery semantics differ. Skipping the load balancer means the timeout behaviour in REL 5.1 is never exercised.

Where a substitution is genuinely necessary, record it and record what it therefore does not test. An explicit list of untested behaviours is useful. An implicit one is discovered during an incident.

Be honest about the environments that exist rather than the ones on the diagram. Most organizations have more than they think, and the ones nobody maintains are usually the ones people test against.

On STACKIT. Managed services help here, because the same service at a smaller plan is still the same service. A PostgreSQL Flex replica set at a small size behaves like a replica set; a single instance does not, whatever its size. Choosing a smaller plan of the correct topology preserves shape, and it is the cheaper of the two ways to save money on a pre-production environment.

The same applies to Kubernetes Engine : a multi-zone node pool with fewer nodes tests the scheduling and storage anchoring behaviour that REL 4.2 describes, while a single-zone pool does not.

Tradeoffs. Cost Optimization. Consistent shape costs more than a minimal stand-in, and this is the most common reason environments diverge. The counter-argument is that a pre-production environment which does not predict production is spending money for no information.

Verify. List the ways your staging environment differs from production. For each, what behaviour is therefore not tested before a change reaches users?


OPS 6.2 Build every environment from the same definitions

Section titled “OPS 6.2 Build every environment from the same definitions”

Risk if not established: High

If environments are built from the same code with different parameters, they cannot silently diverge in shape. If each has its own definition, divergence is guaranteed, because changes will be applied to one and not the others.

Parameterize what should differ, which is sizes, counts and endpoint names, and share everything else. The test of whether you have the split right is whether adding a component to production requires touching anything other than a parameter file.

Environment-specific exceptions are where this erodes. Each one is individually reasonable and collectively they reconstruct the separate definitions you were avoiding. Treat an exception as something that needs a reason and a review, in the same way OPS 3.2 treats an infrastructure change.

The strongest version is being able to create a new environment from scratch on demand. That capability is what makes the rehearsals in REL 8.3 and REL 9.3 affordable, and it is the same mechanism.

On STACKIT. This is OPS 3 applied per environment: a Terraform module or Pulumi program with a variable file per environment, rather than a copy per environment.

The Resource Manager hierarchy is the natural boundary: one project per environment gives separate access, separate quotas and separate billing while the definitions stay shared. Separate projects are also what make SEC 2 and COST 2 work, so the same structure serves three purposes.

Tradeoffs. Operational Excellence. Shared definitions mean a change intended for one environment can affect all of them if the parameterization is wrong, which is what the review in OPS 3.2 catches. Cost Optimization. Creating full environments on demand consumes resources while they exist, which is the argument for destroying them afterwards under SUS 6.

Verify. Are your environments created from the same definitions with different parameters? How many environment-specific exceptions exist, and what is each one for?


OPS 6.3 Keep platform versions aligned, and upgrade in the same order as you deploy

Section titled “OPS 6.3 Keep platform versions aligned, and upgrade in the same order as you deploy”

Risk if not established: Medium

Version drift is the most common shape difference and the least visible, because nothing looks different until a behaviour changes. A Kubernetes minor version, a database major version or a runtime version that differs between environments means the earlier stage is testing against something other than what production runs.

Upgrade in the same order as you deploy: pre-production first, soak, then production. That turns the upgrade itself into a change that passes through your deployment stages rather than a separate class of work that bypasses them.

Bound the acceptable gap. A pre-production environment permanently one version ahead is a useful early warning system; one that is three versions ahead is a different system. Decide the gap deliberately and treat exceeding it as drift under OPS 3.3.

Managed service upgrades happen on the provider’s schedule as well as yours, which means alignment is something to maintain rather than something to set once. Each managed service publishes a lifecycle reference with dated end-of-support per version, such as PostgreSQL Flex , which is what lets an upgrade be planned rather than triggered.

On STACKIT. The cluster version lifecycle and the maintenance window are both configuration, and keeping them in code per OPS 3.1 is what stops environments drifting apart through separate manual settings.

Aus der STACKIT-DokuVersion updates › Maintenance window behaviorStand der Quelle 17.09.2026 · übernommen 05.10.2026

The maintenance window defines when SKE starts maintenance operations such as node rolling updates.

SKE can not guarantee that a started maintenance operation will finish within the configured maintenance window. For large clusters, rolling updates can take several hours. The total duration depends on your workload, your configured Pod Disruption Budgets, and termination grace periods.

Internally, SKE uses a 15-minute buffer and tries to finish maintenance operations before the configured end of the maintenance window. This is a best-effort behavior and can still exceed the window.

If you do not configure a maintenance window, SKE computes one automatically. You can change the maintenance window later.

Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

For Compute Engine, Server Update Management provides scheduled operating system updates for Linux and Windows, and update schedules are where the ordering between environments is expressed. Scheduling pre-production ahead of production is a small configuration choice that turns patching into a tested change.

Tradeoffs. Reliability. Upgrading pre-production first means it is the environment most likely to break, which is the point and is still disruptive to whoever is using it.

Verify. For each managed service, which version runs in production and which in pre-production? How is that gap bounded, and who notices when it grows?


OPS 6.4 Test with production-like data without using production data

Section titled “OPS 6.4 Test with production-like data without using production data”

Risk if not established: High

Data shape drives behaviour more than most teams expect. Query plans change with volume and distribution, edge cases live in the long tail, and a test data set of a thousand tidy rows exercises none of the paths that a production data set with skew, nulls and legacy records will.

Copying production data solves the realism problem and creates a much larger one. The copy carries the same classification, the same regulatory obligations and the same breach consequence, in an environment with weaker controls and broader access. That is a finding under SEC 3 and SOV 3 regardless of how useful it is.

The workable middle is generated or transformed data that preserves the properties that matter: volume, cardinality, distribution and the awkward shapes, without the identifying content. Generating it well is genuine work and it is a one-time cost against a permanent risk.

Where regulated data is involved the answer is simpler, because it is not a judgement call. Under SOV 1, a Tier 1 data set does not get copied into a test environment.

On STACKIT. Instance cloning in PostgreSQL Flex makes copying easy, which is useful for a recovery rehearsal under REL 8.3 and is exactly the mechanism to be careful with here. A clone of production is production data.

Where a copy is genuinely justified, the controls travel with it: encryption under SEC 7, restricted access under SEC 5, residency and retention under SOV 3, and deletion when the test is finished rather than when someone remembers.

Tradeoffs. Performance Efficiency. Realistic data volumes in pre-production cost storage and make environments slower to create. Security. Any production copy is an expansion of the attack surface, which is why the generated alternative is worth the effort.

Verify. Where does your pre-production data come from? If it is a copy of production, which controls travelled with it, and when is it deleted?


  • OPS 3 Everything as code, which supplies the shared definitions
  • OPS 5 Safe deployment, whose earlier stages only predict production if this holds
  • REL 8.3 and REL 9.3, whose rehearsals need environments created on demand
  • SEC 3 Data classification and SOV 1 Sovereignty tier, which govern test data
  • SUS 6 Shutting down idle, which applies to environments created on demand