---
id: COST03
pillar: cost-optimization
title: COST 3. How do you keep provisioned capacity matched to measured demand?
description: A right-sized system drifts out of true without anyone doing anything wrong. Sizing is therefore a cadence rather than a decision made once at the start.
status: draft
services: [compute-engine, postgresql-flex]
sidebar:
  order: 12
  label: Right-sizing
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/cost-03-right-sizing/"
source_file: "docs/architecture/pillars/cost-optimization/cost-03-right-sizing.mdx"
---

Sizing decays. Traffic grows, features change the access pattern, someone adds an index, a
dependency gets faster. None of this is anyone's fault and all of it means the size chosen six
months ago is now approximately wrong, usually in the expensive direction.

This is the same exercise as [`PERF 3.1`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-31-size-from-measured-demand-rather-than-from-a-starting-estimate) seen from the other side. Performance asks whether there
is enough; cost asks whether there is too much. They agree more often than not, and where they
disagree it is because performance wants headroom and cost wants utilization.

## Best practices

- [`COST 3.1`](/architecture/pillars/cost-optimization/cost-03-right-sizing/#cost-31-compare-provisioned-capacity-against-measured-usage-on-a-cadence) Compare provisioned capacity against measured usage on a cadence
- [`COST 3.2`](/architecture/pillars/cost-optimization/cost-03-right-sizing/#cost-32-treat-idle-resources-as-defects-rather-than-as-slack) Treat idle resources as defects rather than as slack
- [`COST 3.3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/#cost-33-know-which-resizings-are-cheap-and-which-are-migrations) Know which resizings are cheap and which are migrations
- [`COST 3.4`](/architecture/pillars/cost-optimization/cost-03-right-sizing/#cost-34-right-size-the-shape-not-only-the-size) Right-size the shape, not only the size

---

## COST 3.1 Compare provisioned capacity against measured usage on a cadence

**Risk if not established:** Medium

Without a cadence, sizing is revisited when a budget review forces it, which is the most expensive
moment to do it and the one where the decisions are worst.

Compare the two numbers per component: what is provisioned and what is used at peak. The gap is
the opportunity, and it is usually larger than expected because sizing decisions are made once
with a safety margin and then inherited.

Measure the peak rather than the average, at a resolution that catches bursts. [`PERF 3.1`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-31-size-from-measured-demand-rather-than-from-a-starting-estimate) makes
the same point for the opposite reason: an average that looks comfortable can hide a saturation,
and a peak that looks alarming can be a five-second burst that nothing depends on.

Do it alongside the performance review under [`PERF 9.1`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/#perf-91-re-examine-sizing-against-measured-demand-on-a-cadence). The two look at the same measurements and
reach conclusions that need reconciling, and reconciling them in one conversation is cheaper than
discovering the conflict afterwards.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost
Dashboard</LinkChip> supplies the
spend side per project, with monthly, quarterly, half-yearly, yearly and user-defined ranges. The
usage side comes from
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> and
your instrumentation under [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/). Neither is useful alone: spend without utilization tells you
what you pay and not whether it is warranted.

**Tradeoffs.** **Operational Excellence.** A recurring review costs time from people who could be
building, and its output is frequently a change that carries its own risk.

**Verify.** For your five largest cost items, what is provisioned and what is used at peak? When
was that comparison last made?

---

## COST 3.2 Treat idle resources as defects rather than as slack

**Risk if not established:** Medium

There is a difference between headroom and waste, and it is whether anyone decided. Headroom is
capacity reserved for a stated reason under [`REL 7.1`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-71-size-for-measured-peaks-and-keep-headroom-for-the-peak-you-did-not-predict). Waste is capacity nobody is aware of.

The recurring finds: volumes detached from any instance, instances stopped but still billing, load
balancers with no backends, snapshots from a migration that finished, environments from a project
that ended, and IP addresses reserved and unused.

Each is small and they accumulate, and none of them will ever be found by looking at the largest
line items, which is why [`COST 7.2`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/#cost-72-look-for-the-spend-that-buys-least-rather-than-the-largest-line-item) argues for looking at the spend that buys least rather than
the spend that is biggest.

Make finding them recurring rather than heroic. A quarterly sweep that deletes twenty forgotten
resources is worth more than an annual project, because the resources were created continuously.

**On STACKIT.** One billing behaviour is the most common source of an instance that costs money
while doing nothing: a machine that is merely stopped
keeps its resource reservation and continues to be billed for it, while a shelved machine does
not. Both look like the instance is off. [`COST 5.2`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-52-shut-down-what-is-not-being-used-and-know-what-stopping-actually-stops) carries the distinction in full, with the
service-certificate wording it rests on.

**Tradeoffs.** **Reliability.** Deleting something that turns out to be needed is the risk, which
is why the ownership from [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims) comes first. An unclaimed resource is easier to delete
confidently than an unlabelled one.

**Verify.** How many volumes in your estate are not attached to anything? How many instances are
stopped rather than shelved?

---

## COST 3.3 Know which resizings are cheap and which are migrations

**Risk if not established:** Medium

Right-sizing assumes the size can be changed. Where it cannot, the decision is not whether to
resize but whether to migrate, and that changes the arithmetic entirely.

An over-provisioned component whose size is a setting should be corrected as soon as it is
noticed. One whose correction requires a migration needs the saving weighed against the effort and
the risk, and the answer is sometimes to leave it.

This is [`PERF 3.2`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-32-establish-which-sizing-decisions-are-reversible-before-you-make-them) from the cost side and the sort is the same: a setting, a replacement, or a
migration. Making it once and recording it serves both pillars.

The asymmetry is worth naming. Sizing up under pressure and sizing down at leisure have very
different urgency, so an irreversible axis should be sized with more margin than a reversible one,
which means accepting some waste deliberately.

**On STACKIT.** Compute Engine offers <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server/basics/machine-types/">machine
types</LinkChip> as fixed
variants, and a configuration cannot be adapted beyond them.
Changing type is a replacement of the instance, which is disruptive and bounded.

Managed database <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/">performance
classes</LinkChip>
are the case where a size correction is a migration: a class that is too small requires cloning to
a new instance. Downsizing an over-provisioned class carries the
same cost, which means an over-provisioned database is frequently cheaper to leave than to
correct.

**Tradeoffs.** **Cost Optimization**, against itself. Accepting known waste on an irreversible
axis is a deliberate cost, justified by the risk and effort of the alternative.

**Verify.** For each over-provisioned component, is correcting it a setting, a replacement or a
migration? For the migrations, is the saving worth the work?

---

## COST 3.4 Right-size the shape, not only the size

**Risk if not established:** Medium

Sizing on one dimension means over-provisioning every other dimension to obtain enough of the one
that binds. A memory-bound service on a CPU-weighted instance pays for cores it does not use.

Establish which resource actually limits the component, which is the measurement [`PERF 3.3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-33-match-the-resource-shape-to-the-workload-shape) and
[`PERF 7.1`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/#perf-71-profile-to-find-the-constraint-before-changing-anything) both need, then choose an option weighted toward it. The saving comes from no longer
buying the dimensions that were never the constraint.

Storage tiers are the dimension most often ignored, because they are not visible in the vCPU and
RAM figures that dominate a sizing conversation. A workload on a high I/O tier that performs few
operations is paying for throughput it does not consume.

Watch for components whose shape changed. A service that was CPU-bound before a cache was added
may be memory-bound after it, and the instance chosen for the old shape is now wrong in a new
direction.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server/basics/machine-types/">machine type
families</LinkChip> are
organized on exactly this axis, by vCPU-to-RAM ratio from CPU-weighted through general purpose to
memory-weighted. Choosing the family is choosing the shape and matters more than choosing the size
within one.

Managed databases separate the two dimensions: the flavor covers compute and memory, the
performance class covers I/O, and they are chosen independently. That is useful and means both can
be over-provisioned independently, so both belong in the review.

**Tradeoffs.** None material. Matching the shape reduces cost and improves performance at once,
which is one of the few places where the two pillars agree without qualification.

**Verify.** For your largest component, which resource is the binding constraint? Does its shape
weight that resource, or was it chosen on total size?

---

## Related

- [`PERF 3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/) Selection and sizing, the same decisions from the performance side
- [`PERF 9.1`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/#perf-91-re-examine-sizing-against-measured-demand-on-a-cadence) Lifecycle review, which should happen in the same conversation
- [`COST 5`](/architecture/pillars/cost-optimization/cost-05-environments/) Environments, where the largest idle capacity usually sits
- [`REL 7.1`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-71-size-for-measured-peaks-and-keep-headroom-for-the-peak-you-did-not-predict) Headroom, which is the deliberate version of unused capacity
- [`SUS 2`](/architecture/pillars/sustainability/sus-02-right-sizing/) Right-sizing, the same lever with a different stopping point
