Zum Inhalt springen
Beta

COST 3. How do you keep provisioned capacity matched to measured demand?

Zuletzt aktualisiert am

Sizing decays. Traffic grows, features change the access pattern, someone adds an index, a dependency gets faster. None of this is anyone’s fault and all of it means the size chosen six months ago is now approximately wrong, usually in the expensive direction.

This is the same exercise as PERF 3.1 seen from the other side. Performance asks whether there is enough; cost asks whether there is too much. They agree more often than not, and where they disagree it is because performance wants headroom and cost wants utilization.

  • COST 3.1 Compare provisioned capacity against measured usage on a cadence
  • COST 3.2 Treat idle resources as defects rather than as slack
  • COST 3.3 Know which resizings are cheap and which are migrations
  • COST 3.4 Right-size the shape, not only the size

COST 3.1 Compare provisioned capacity against measured usage on a cadence

Section titled “COST 3.1 Compare provisioned capacity against measured usage on a cadence”

Risk if not established: Medium

Without a cadence, sizing is revisited when a budget review forces it, which is the most expensive moment to do it and the one where the decisions are worst.

Compare the two numbers per component: what is provisioned and what is used at peak. The gap is the opportunity, and it is usually larger than expected because sizing decisions are made once with a safety margin and then inherited.

Measure the peak rather than the average, at a resolution that catches bursts. PERF 3.1 makes the same point for the opposite reason: an average that looks comfortable can hide a saturation, and a peak that looks alarming can be a five-second burst that nothing depends on.

Do it alongside the performance review under PERF 9.1. The two look at the same measurements and reach conclusions that need reconciling, and reconciling them in one conversation is cheaper than discovering the conflict afterwards.

On STACKIT. The Cost Dashboard supplies the spend side per project, with monthly, quarterly, half-yearly, yearly and user-defined ranges. The usage side comes from Observability and your instrumentation under OPS 7. Neither is useful alone: spend without utilization tells you what you pay and not whether it is warranted.

Tradeoffs. Operational Excellence. A recurring review costs time from people who could be building, and its output is frequently a change that carries its own risk.

Verify. For your five largest cost items, what is provisioned and what is used at peak? When was that comparison last made?


COST 3.2 Treat idle resources as defects rather than as slack

Section titled “COST 3.2 Treat idle resources as defects rather than as slack”

Risk if not established: Medium

There is a difference between headroom and waste, and it is whether anyone decided. Headroom is capacity reserved for a stated reason under REL 7.1. Waste is capacity nobody is aware of.

The recurring finds: volumes detached from any instance, instances stopped but still billing, load balancers with no backends, snapshots from a migration that finished, environments from a project that ended, and IP addresses reserved and unused.

Each is small and they accumulate, and none of them will ever be found by looking at the largest line items, which is why COST 7.2 argues for looking at the spend that buys least rather than the spend that is biggest.

Make finding them recurring rather than heroic. A quarterly sweep that deletes twenty forgotten resources is worth more than an annual project, because the resources were created continuously.

On STACKIT. One billing behaviour is the most common source of an instance that costs money while doing nothing: a machine that is merely stopped keeps its resource reservation and continues to be billed for it, while a shelved machine does not. Both look like the instance is off. COST 5.2 carries the distinction in full, with the service-certificate wording it rests on.

Tradeoffs. Reliability. Deleting something that turns out to be needed is the risk, which is why the ownership from COST 2.4 comes first. An unclaimed resource is easier to delete confidently than an unlabelled one.

Verify. How many volumes in your estate are not attached to anything? How many instances are stopped rather than shelved?


COST 3.3 Know which resizings are cheap and which are migrations

Section titled “COST 3.3 Know which resizings are cheap and which are migrations”

Risk if not established: Medium

Right-sizing assumes the size can be changed. Where it cannot, the decision is not whether to resize but whether to migrate, and that changes the arithmetic entirely.

An over-provisioned component whose size is a setting should be corrected as soon as it is noticed. One whose correction requires a migration needs the saving weighed against the effort and the risk, and the answer is sometimes to leave it.

This is PERF 3.2 from the cost side and the sort is the same: a setting, a replacement, or a migration. Making it once and recording it serves both pillars.

The asymmetry is worth naming. Sizing up under pressure and sizing down at leisure have very different urgency, so an irreversible axis should be sized with more margin than a reversible one, which means accepting some waste deliberately.

On STACKIT. Compute Engine offers machine types as fixed variants, and a configuration cannot be adapted beyond them. Changing type is a replacement of the instance, which is disruptive and bounded.

Managed database performance classes are the case where a size correction is a migration: a class that is too small requires cloning to a new instance. Downsizing an over-provisioned class carries the same cost, which means an over-provisioned database is frequently cheaper to leave than to correct.

Tradeoffs. Cost Optimization, against itself. Accepting known waste on an irreversible axis is a deliberate cost, justified by the risk and effort of the alternative.

Verify. For each over-provisioned component, is correcting it a setting, a replacement or a migration? For the migrations, is the saving worth the work?


COST 3.4 Right-size the shape, not only the size

Section titled “COST 3.4 Right-size the shape, not only the size”

Risk if not established: Medium

Sizing on one dimension means over-provisioning every other dimension to obtain enough of the one that binds. A memory-bound service on a CPU-weighted instance pays for cores it does not use.

Establish which resource actually limits the component, which is the measurement PERF 3.3 and PERF 7.1 both need, then choose an option weighted toward it. The saving comes from no longer buying the dimensions that were never the constraint.

Storage tiers are the dimension most often ignored, because they are not visible in the vCPU and RAM figures that dominate a sizing conversation. A workload on a high I/O tier that performs few operations is paying for throughput it does not consume.

Watch for components whose shape changed. A service that was CPU-bound before a cache was added may be memory-bound after it, and the instance chosen for the old shape is now wrong in a new direction.

On STACKIT. The machine type families are organized on exactly this axis, by vCPU-to-RAM ratio from CPU-weighted through general purpose to memory-weighted. Choosing the family is choosing the shape and matters more than choosing the size within one.

Managed databases separate the two dimensions: the flavor covers compute and memory, the performance class covers I/O, and they are chosen independently. That is useful and means both can be over-provisioned independently, so both belong in the review.

Tradeoffs. None material. Matching the shape reduces cost and improves performance at once, which is one of the few places where the two pillars agree without qualification.

Verify. For your largest component, which resource is the binding constraint? Does its shape weight that resource, or was it chosen on total size?


  • PERF 3 Selection and sizing, the same decisions from the performance side
  • PERF 9.1 Lifecycle review, which should happen in the same conversation
  • COST 5 Environments, where the largest idle capacity usually sits
  • REL 7.1 Headroom, which is the deliberate version of unused capacity
  • SUS 2 Right-sizing, the same lever with a different stopping point