OPS 6. How do you keep environments consistent from development to production?
Zuletzt aktualisiert am
Every other question in this pillar assumes that what happens in a pre-production environment predicts what will happen in production. When environments diverge, that assumption quietly stops holding, and the tests, the deployment stages and the rehearsals all keep running while telling you less and less.
Divergence is never decided. It accumulates: a version upgraded in one place, a component simplified because the full one was expensive, a configuration mechanism that differs because someone was in a hurry.
Best practices
Section titled “Best practices”OPS 6.1Let environments differ in scale and data, not in shapeOPS 6.2Build every environment from the same definitionsOPS 6.3Keep platform versions aligned, and upgrade in the same order as you deployOPS 6.4Test with production-like data without using production data
OPS 6.1 Let environments differ in scale and data, not in shape
Section titled “OPS 6.1 Let environments differ in scale and data, not in shape”Risk if not established: High
Shape is topology, configuration mechanism, component set and versions. Scale is instance sizes, replica counts and data volume. Differences in scale are expected and manageable. Differences in shape are what make a pre-production result meaningless.
The substitutions that cause the most trouble are the ones that feel harmless. Running a single
database instance instead of a replica set means the failover behaviour is untested. Using an
in-memory queue instead of the real broker means the delivery semantics differ. Skipping the load
balancer means the timeout behaviour in REL 5.1 is never exercised.
Where a substitution is genuinely necessary, record it and record what it therefore does not test. An explicit list of untested behaviours is useful. An implicit one is discovered during an incident.
Be honest about the environments that exist rather than the ones on the diagram. Most organizations have more than they think, and the ones nobody maintains are usually the ones people test against.
On STACKIT. Managed services help here, because the same service at a smaller plan is still the same service. A PostgreSQL Flex replica set at a small size behaves like a replica set; a single instance does not, whatever its size. Choosing a smaller plan of the correct topology preserves shape, and it is the cheaper of the two ways to save money on a pre-production environment.
The same applies to Kubernetes
Engine : a
multi-zone node pool with fewer nodes tests the scheduling and storage anchoring behaviour that
REL 4.2 describes, while a single-zone pool does not.
Tradeoffs. Cost Optimization. Consistent shape costs more than a minimal stand-in, and this is the most common reason environments diverge. The counter-argument is that a pre-production environment which does not predict production is spending money for no information.
Verify. List the ways your staging environment differs from production. For each, what behaviour is therefore not tested before a change reaches users?
OPS 6.2 Build every environment from the same definitions
Section titled “OPS 6.2 Build every environment from the same definitions”Risk if not established: High
If environments are built from the same code with different parameters, they cannot silently diverge in shape. If each has its own definition, divergence is guaranteed, because changes will be applied to one and not the others.
Parameterize what should differ, which is sizes, counts and endpoint names, and share everything else. The test of whether you have the split right is whether adding a component to production requires touching anything other than a parameter file.
Environment-specific exceptions are where this erodes. Each one is individually reasonable and
collectively they reconstruct the separate definitions you were avoiding. Treat an exception as
something that needs a reason and a review, in the same way OPS 3.2 treats an infrastructure
change.
The strongest version is being able to create a new environment from scratch on demand. That
capability is what makes the rehearsals in REL 8.3 and REL 9.3 affordable, and it is the same
mechanism.
On STACKIT. This is OPS 3 applied per environment: a
Terraform
module or Pulumi program with a
variable file per environment, rather than a copy per environment.
The Resource Manager
hierarchy is the natural boundary:
one project per environment gives separate access, separate quotas and separate billing while the
definitions stay shared. Separate projects are also what make SEC 2 and COST 2 work, so the
same structure serves three purposes.
Tradeoffs. Operational Excellence. Shared definitions mean a change intended for one
environment can affect all of them if the parameterization is wrong, which is what the review in
OPS 3.2 catches. Cost Optimization. Creating full environments on demand consumes resources
while they exist, which is the argument for destroying them afterwards under SUS 6.
Verify. Are your environments created from the same definitions with different parameters? How many environment-specific exceptions exist, and what is each one for?
OPS 6.3 Keep platform versions aligned, and upgrade in the same order as you deploy
Section titled “OPS 6.3 Keep platform versions aligned, and upgrade in the same order as you deploy”Risk if not established: Medium
Version drift is the most common shape difference and the least visible, because nothing looks different until a behaviour changes. A Kubernetes minor version, a database major version or a runtime version that differs between environments means the earlier stage is testing against something other than what production runs.
Upgrade in the same order as you deploy: pre-production first, soak, then production. That turns the upgrade itself into a change that passes through your deployment stages rather than a separate class of work that bypasses them.
Bound the acceptable gap. A pre-production environment permanently one version ahead is a useful
early warning system; one that is three versions ahead is a different system. Decide the gap
deliberately and treat exceeding it as drift under OPS 3.3.
Managed service upgrades happen on the provider’s schedule as well as yours, which means alignment is something to maintain rather than something to set once. Each managed service publishes a lifecycle reference with dated end-of-support per version, such as PostgreSQL Flex , which is what lets an upgrade be planned rather than triggered.
On STACKIT. The cluster version
lifecycle
and the maintenance
window
are both configuration, and keeping them in code per OPS 3.1 is what stops environments drifting
apart through separate manual settings.
The maintenance window defines when SKE starts maintenance operations such as node rolling updates.
SKE can not guarantee that a started maintenance operation will finish within the configured maintenance window. For large clusters, rolling updates can take several hours. The total duration depends on your workload, your configured Pod Disruption Budgets, and termination grace periods.
Internally, SKE uses a 15-minute buffer and tries to finish maintenance operations before the configured end of the maintenance window. This is a best-effort behavior and can still exceed the window.
If you do not configure a maintenance window, SKE computes one automatically. You can change the maintenance window later.
Was ist das?
Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.
For Compute Engine, Server Update Management provides scheduled operating system updates for Linux and Windows, and update schedules are where the ordering between environments is expressed. Scheduling pre-production ahead of production is a small configuration choice that turns patching into a tested change.
Tradeoffs. Reliability. Upgrading pre-production first means it is the environment most likely to break, which is the point and is still disruptive to whoever is using it.
Verify. For each managed service, which version runs in production and which in pre-production? How is that gap bounded, and who notices when it grows?
OPS 6.4 Test with production-like data without using production data
Section titled “OPS 6.4 Test with production-like data without using production data”Risk if not established: High
Data shape drives behaviour more than most teams expect. Query plans change with volume and distribution, edge cases live in the long tail, and a test data set of a thousand tidy rows exercises none of the paths that a production data set with skew, nulls and legacy records will.
Copying production data solves the realism problem and creates a much larger one. The copy carries
the same classification, the same regulatory obligations and the same breach consequence, in an
environment with weaker controls and broader access. That is a finding under SEC 3 and SOV 3
regardless of how useful it is.
The workable middle is generated or transformed data that preserves the properties that matter: volume, cardinality, distribution and the awkward shapes, without the identifying content. Generating it well is genuine work and it is a one-time cost against a permanent risk.
Where regulated data is involved the answer is simpler, because it is not a judgement call. Under
SOV 1, a Tier 1 data set does not get copied into a test environment.
On STACKIT. Instance cloning in PostgreSQL
Flex
makes copying easy, which is useful for a recovery rehearsal under REL 8.3 and is exactly the
mechanism to be careful with here. A clone of production is production data.
Where a copy is genuinely justified, the controls travel with it: encryption under SEC 7,
restricted access under SEC 5, residency and retention under SOV 3, and deletion when the test
is finished rather than when someone remembers.
Tradeoffs. Performance Efficiency. Realistic data volumes in pre-production cost storage and make environments slower to create. Security. Any production copy is an expansion of the attack surface, which is why the generated alternative is worth the effort.
Verify. Where does your pre-production data come from? If it is a copy of production, which controls travelled with it, and when is it deleted?
Related
Section titled “Related”OPS 3Everything as code, which supplies the shared definitionsOPS 5Safe deployment, whose earlier stages only predict production if this holdsREL 8.3andREL 9.3, whose rehearsals need environments created on demandSEC 3Data classification andSOV 1Sovereignty tier, which govern test dataSUS 6Shutting down idle, which applies to environments created on demand