Zum Inhalt springen
Beta

COST 5. How do you size non-production environments for their real purpose?

Zuletzt aktualisiert am

Non-production environments are created for a reason, sized generously so they do not get in the way, and then never revisited. They run continuously, including the two thirds of the week when nobody is working, and nobody receives a bill with their name on it.

This is usually the largest available saving in an estate and the least contentious, because reducing it costs nobody anything they were using.

  • COST 5.1 Size each environment for what it is actually used for
  • COST 5.2 Shut down what is not being used, and know what stopping actually stops
  • COST 5.3 Preserve the shape where testing depends on it, reduce the scale
  • COST 5.4 Create environments on demand rather than keeping them

COST 5.1 Size each environment for what it is actually used for

Section titled “COST 5.1 Size each environment for what it is actually used for”

Risk if not established: Medium

Environments get sized by copying production and reducing it a little, which produces an environment sized for a purpose it does not have.

Establish the actual purpose first. An environment for functional testing needs correctness rather than capacity. One for integration testing needs the real components at small scale. One for performance testing needs production-like shape and enough scale for the result to transfer, per PERF 8.3. One for demonstrations needs to look right for an hour a month.

Those are four different sizes, and treating them as one is where the money goes.

Enumerate them honestly. Most organizations have more environments than the diagram shows, and the ones nobody maintains are frequently the ones still running. SEC 1.4 and COST 2.4 find the same list.

On STACKIT. Separate projects per environment make each one’s cost visible in the Cost Dashboard without any extra work, since it breaks down per project. Environments sharing a project have combined costs that no later analysis can separate, which is the same argument COST 2.1 makes.

Tradeoffs. Operational Excellence. Differently sized environments diverge from production in ways that matter, which is OPS 6.1. The resolution is in COST 5.3: reduce the scale, keep the shape.

Verify. List your environments and what each costs per month. For each, what is it actually used for, and how many hours per week is it used?


COST 5.2 Shut down what is not being used, and know what stopping actually stops

Section titled “COST 5.2 Shut down what is not being used, and know what stopping actually stops”

Risk if not established: Medium

An environment used during working hours runs for about a quarter of the week and is billed for all of it. Shutting it down outside those hours removes most of its cost without removing anything anyone uses.

Automate it rather than relying on someone remembering, which is OPS 10.2. A schedule that stops things in the evening and starts them in the morning is a small piece of automation with a return that recurs every day.

The part that catches people out is that stopping is not always the same as not being billed. Compute resources are frequently billed on reservation rather than on execution, and a machine that is “off” may still be holding the resources it reserved.

Check per resource type. Storage almost always continues to be billed regardless of whether anything is running, so an environment that is stopped nightly still pays for its volumes, and that residual is the floor of what shutting down can save.

On STACKIT. This distinction is documented precisely, and it is the most useful billing fact in this pillar. The Compute Engine service certificate states that the billed period runs from creation to deletion minus any shelving periods, and defines shelving as stopping the machine with its resource reservation cancelled.

So a machine that is merely stopped keeps its reservation and continues to be billed for it, while a shelved machine does not. Both look like the instance is off. An automated shutdown that stops without shelving therefore saves nothing, which is a disappointing thing to discover after building it.

Attached storage is separate from the machine’s reservation, so it continues regardless. That is the residual cost of a shelved environment and the reason COST 5.4 is the stronger answer where it is achievable.

Tradeoffs. Operational Excellence. Start-up time before the environment is usable, and the automation itself is a thing to maintain. Reliability, mildly: an environment that is started on demand is one more thing that can fail to start when someone needs it.

Verify. For each non-production environment, how many hours per week does it run and how many does anyone use it? Of the resources that are stopped outside those hours, which are still billed?


COST 5.3 Preserve the shape where testing depends on it, reduce the scale

Section titled “COST 5.3 Preserve the shape where testing depends on it, reduce the scale”

Risk if not established: Medium

The tension in this question is with OPS 6.1, which requires environments to differ in scale and data rather than in shape. Cutting cost by simplifying the topology is exactly what that best practice forbids, and for good reason: an environment that does not resemble production stops predicting it.

The resolution is to cut on the axis that does not carry the information. A smaller instance of the same service preserves the behaviour. A single instance replacing a replica set does not, because the failover behaviour was the thing being tested.

Decide per environment which properties have to hold. A functional test environment can substitute aggressively; a pre-production environment that gates deployment under OPS 5.4 cannot, because its whole purpose is to predict what production will do.

Where a substitution is necessary, record what it therefore does not test. OPS 6.1 makes the same point and it is worth writing down once for both purposes.

On STACKIT. Managed services help here, because a smaller plan of the same service is still the same service. A PostgreSQL Flex replica set at a small flavor behaves like a replica set; a single instance does not, whatever its size. Choosing a smaller flavor of the correct topology is the cheaper of the two ways to save money, and the only one that preserves the test.

The same applies to Kubernetes Engine : a multi-zone node pool with fewer nodes exercises the scheduling and storage anchoring behaviour that REL 4.2 describes, while a single-zone pool does not.

Tradeoffs. Operational Excellence, directly. Every saving on this axis is a reduction in what the environment tells you, which is why the decision is per environment rather than uniform.

Verify. For each non-production environment, list the ways it differs from production. Which of those differences are scale, and which are shape?


COST 5.4 Create environments on demand rather than keeping them

Section titled “COST 5.4 Create environments on demand rather than keeping them”

Risk if not established: Medium

An environment that exists only while it is needed costs only while it is needed, which is a stronger result than any amount of scheduling around a permanent one.

It requires the environment to be creatable from definitions, which is OPS 3 and OPS 6.2. Where that capability exists it pays several times over: for the environment cost here, for the recovery rehearsals in REL 8.3 and REL 9.3, and for the load testing in PERF 8.3.

Not everything suits it. An environment holding long-lived state that is expensive to reproduce, or one that takes hours to become usable, is better kept and scheduled. The candidates are the ones created for a purpose with a beginning and an end: a feature branch, a load test, a migration rehearsal, a demonstration.

Destroy them by default rather than on request. An environment created on demand and never destroyed is a permanent environment that nobody planned, and those are the ones that end up in the unattributed remainder from COST 2.4.

On STACKIT. The Terraform provider , Pulumi , the CLI and the API are what make creation and destruction a pipeline step rather than a project.

Creating a project per ephemeral environment keeps its cost separately visible and makes destruction clean, and the 2,500 project limit per organization is high enough that this is viable. One constraint to design around: a project is linked to one billing account and reassignment is not currently possible, so an ephemeral project inherits whatever billing arrangement it was created under.

Tradeoffs. Operational Excellence. On-demand creation is a capability to build and maintain, and it only pays where it is used often enough. Sustainability. It is the strongest version of SUS 6, since a destroyed environment consumes nothing at all.

Verify. Which of your environments could be created on demand? For those that exist permanently, what state do they hold that could not be reproduced?


  • OPS 6 Environment consistency, which this question is in direct tension with
  • OPS 3 Everything as code, without which on-demand creation is not available
  • COST 2.1 Attribution, which makes per-environment cost visible
  • SUS 6 Shutting down idle, the same lever with a different motive
  • PERF 8.3 Load testing, which needs an environment the result transfers from