---
id: COST05
pillar: cost-optimization
title: COST 5. How do you size non-production environments for their real purpose?
description: "Non-production is where the quiet waste accumulates: nobody watches it, nobody is billed for it, and it runs through nights and weekends serving nobody."
status: draft
services: [compute-engine, kubernetes-engine]
sidebar:
  order: 14
  label: Environments
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/cost-05-environments/"
source_file: "docs/architecture/pillars/cost-optimization/cost-05-environments.mdx"
---

Non-production environments are created for a reason, sized generously so they do not get in the
way, and then never revisited. They run continuously, including the two thirds of the week when
nobody is working, and nobody receives a bill with their name on it.

This is usually the largest available saving in an estate and the least contentious, because
reducing it costs nobody anything they were using.

## Best practices

- [`COST 5.1`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-51-size-each-environment-for-what-it-is-actually-used-for) Size each environment for what it is actually used for
- [`COST 5.2`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-52-shut-down-what-is-not-being-used-and-know-what-stopping-actually-stops) Shut down what is not being used, and know what stopping actually stops
- [`COST 5.3`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-53-preserve-the-shape-where-testing-depends-on-it-reduce-the-scale) Preserve the shape where testing depends on it, reduce the scale
- [`COST 5.4`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-54-create-environments-on-demand-rather-than-keeping-them) Create environments on demand rather than keeping them

---

## COST 5.1 Size each environment for what it is actually used for

**Risk if not established:** Medium

Environments get sized by copying production and reducing it a little, which produces an
environment sized for a purpose it does not have.

Establish the actual purpose first. An environment for functional testing needs correctness rather
than capacity. One for integration testing needs the real components at small scale. One for
performance testing needs production-like shape and enough scale for the result to transfer, per
[`PERF 8.3`](/architecture/pillars/performance-efficiency/perf-08-load-testing/#perf-83-test-where-the-result-transfers-to-production). One for demonstrations needs to look right for an hour a month.

Those are four different sizes, and treating them as one is where the money goes.

Enumerate them honestly. Most organizations have more environments than the diagram shows, and the
ones nobody maintains are frequently the ones still running. [`SEC 1.4`](/architecture/pillars/security/sec-01-security-baseline/#sec-14-cover-the-whole-estate-including-the-parts-nobody-claims) and [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims) find the
same list.

**On STACKIT.** Separate projects per environment make each one's cost visible in the
<LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost Dashboard</LinkChip> without any
extra work, since it breaks down per project. Environments sharing a project have combined costs
that no later analysis can separate, which is the same argument [`COST 2.1`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-21-make-the-resource-hierarchy-reflect-who-pays) makes.

**Tradeoffs.** **Operational Excellence.** Differently sized environments diverge from production
in ways that matter, which is [`OPS 6.1`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-61-let-environments-differ-in-scale-and-data-not-in-shape). The resolution is in [`COST 5.3`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-53-preserve-the-shape-where-testing-depends-on-it-reduce-the-scale): reduce the scale, keep
the shape.

**Verify.** List your environments and what each costs per month. For each, what is it actually
used for, and how many hours per week is it used?

---

## COST 5.2 Shut down what is not being used, and know what stopping actually stops

**Risk if not established:** Medium

An environment used during working hours runs for about a quarter of the week and is billed for
all of it. Shutting it down outside those hours removes most of its cost without removing anything
anyone uses.

Automate it rather than relying on someone remembering, which is [`OPS 10.2`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-102-automate-what-recurs-starting-with-the-riskiest-rather-than-the-most-frequent). A schedule that stops
things in the evening and starts them in the morning is a small piece of automation with a return
that recurs every day.

The part that catches people out is that stopping is not always the same as not being billed.
Compute resources are frequently billed on reservation rather than on execution, and a machine
that is "off" may still be holding the resources it reserved.

Check per resource type. Storage almost always continues to be billed regardless of whether
anything is running, so an environment that is stopped nightly still pays for its volumes, and
that residual is the floor of what shutting down can save.

**On STACKIT.** This distinction is documented precisely, and it is the most useful billing fact
in this pillar. The Compute Engine <LinkChip href="https://stackit.com/en/gtc/service-certificates">service
certificate</LinkChip> states that the billed period runs
from creation to deletion **minus any shelving periods**, and defines shelving as stopping the
machine with its resource reservation cancelled.

So a machine that is merely stopped keeps its reservation and continues to be billed for it, while
a shelved machine does not. Both look like the instance is off. An automated shutdown that stops
without shelving therefore saves nothing, which is a disappointing thing to discover after
building it.

Attached storage is separate from the machine's reservation, so it continues regardless. That is
the residual cost of a shelved environment and the reason [`COST 5.4`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-54-create-environments-on-demand-rather-than-keeping-them) is the stronger answer where
it is achievable.

**Tradeoffs.** **Operational Excellence.** Start-up time before the environment is usable, and the
automation itself is a thing to maintain. **Reliability**, mildly: an environment that is started
on demand is one more thing that can fail to start when someone needs it.

**Verify.** For each non-production environment, how many hours per week does it run and how many
does anyone use it? Of the resources that are stopped outside those hours, which are still billed?

---

## COST 5.3 Preserve the shape where testing depends on it, reduce the scale

**Risk if not established:** Medium

The tension in this question is with [`OPS 6.1`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-61-let-environments-differ-in-scale-and-data-not-in-shape), which requires environments to differ in scale and
data rather than in shape. Cutting cost by simplifying the topology is exactly what that best
practice forbids, and for good reason: an environment that does not resemble production stops
predicting it.

The resolution is to cut on the axis that does not carry the information. A smaller instance of
the same service preserves the behaviour. A single instance replacing a replica set does not,
because the failover behaviour was the thing being tested.

Decide per environment which properties have to hold. A functional test environment can substitute
aggressively; a pre-production environment that gates deployment under [`OPS 5.4`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-54-advance-one-environment-at-a-time-and-one-region-at-a-time) cannot, because
its whole purpose is to predict what production will do.

Where a substitution is necessary, record what it therefore does not test. [`OPS 6.1`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-61-let-environments-differ-in-scale-and-data-not-in-shape) makes the
same point and it is worth writing down once for both purposes.

**On STACKIT.** Managed services help here, because a smaller plan of the same service is still
the same service. A <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/basics/plan-your-postgresql-flex-instance/">PostgreSQL
Flex</LinkChip>
replica set at a small flavor behaves like a replica set; a single instance does not, whatever its
size. Choosing a smaller flavor of the correct topology is the cheaper of the two ways to save
money, and the only one that preserves the test.

The same applies to <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/operations/topologies/">Kubernetes
Engine</LinkChip>: a
multi-zone node pool with fewer nodes exercises the scheduling and storage anchoring behaviour
that [`REL 4.2`](/architecture/pillars/reliability/rel-04-redundancy/#rel-42-make-state-redundant-and-know-where-each-data-set-is-anchored) describes, while a single-zone pool does not.

**Tradeoffs.** **Operational Excellence**, directly. Every saving on this axis is a reduction in
what the environment tells you, which is why the decision is per environment rather than uniform.

**Verify.** For each non-production environment, list the ways it differs from production. Which
of those differences are scale, and which are shape?

---

## COST 5.4 Create environments on demand rather than keeping them

**Risk if not established:** Medium

An environment that exists only while it is needed costs only while it is needed, which is a
stronger result than any amount of scheduling around a permanent one.

It requires the environment to be creatable from definitions, which is [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) and [`OPS 6.2`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-62-build-every-environment-from-the-same-definitions).
Where that capability exists it pays several times over: for the environment cost here, for the
recovery rehearsals in [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it) and [`REL 9.3`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-93-rehearse-it-and-time-the-rehearsal-against-the-rto), and for the load testing in [`PERF 8.3`](/architecture/pillars/performance-efficiency/perf-08-load-testing/#perf-83-test-where-the-result-transfers-to-production).

Not everything suits it. An environment holding long-lived state that is expensive to reproduce,
or one that takes hours to become usable, is better kept and scheduled. The candidates are the
ones created for a purpose with a beginning and an end: a feature branch, a load test, a migration
rehearsal, a demonstration.

Destroy them by default rather than on request. An environment created on demand and never
destroyed is a permanent environment that nobody planned, and those are the ones that end up in
the unattributed remainder from [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims).

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-iac/stackit-terraform-provider/">Terraform
provider</LinkChip>,
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-iac/pulumi/">Pulumi</LinkChip>, the
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-cli/">CLI</LinkChip> and the
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-api/">API</LinkChip> are what make creation and
destruction a pipeline step rather than a project.

Creating a project per ephemeral environment keeps its cost separately visible and makes
destruction clean, and the <LinkChip href="https://docs.stackit.cloud/platform/resource-manager/basics/limitations/">2,500 project
limit</LinkChip> per organization
is high enough that this is viable. One constraint to design around: a project is linked to one
billing account and reassignment is not currently possible, so an
ephemeral project inherits whatever billing arrangement it was created under.

**Tradeoffs.** **Operational Excellence.** On-demand creation is a capability to build and
maintain, and it only pays where it is used often enough. **Sustainability.** It is the strongest
version of [`SUS 6`](/architecture/pillars/sustainability/sus-06-shut-down-idle/), since a destroyed environment consumes nothing at all.

**Verify.** Which of your environments could be created on demand? For those that exist
permanently, what state do they hold that could not be reproduced?

---

## Related

- [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) Environment consistency, which this question is in direct tension with
- [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) Everything as code, without which on-demand creation is not available
- [`COST 2.1`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-21-make-the-resource-hierarchy-reflect-who-pays) Attribution, which makes per-environment cost visible
- [`SUS 6`](/architecture/pillars/sustainability/sus-06-shut-down-idle/) Shutting down idle, the same lever with a different motive
- [`PERF 8.3`](/architecture/pillars/performance-efficiency/perf-08-load-testing/#perf-83-test-where-the-result-transfers-to-production) Load testing, which needs an environment the result transfers from
