COST 3. How do you keep provisioned capacity matched to measured demand?
Last updated on
Sizing decays. Traffic grows, features change the access pattern, someone adds an index, a dependency gets faster. None of this is anyone’s fault and all of it means the size chosen six months ago is now approximately wrong, usually in the expensive direction.
This is the same exercise as PERF 3.1 seen from the other side. Performance asks whether there
is enough; cost asks whether there is too much. They agree more often than not, and where they
disagree it is because performance wants headroom and cost wants utilization.
Best practices
Section titled “Best practices”COST 3.1Compare provisioned capacity against measured usage on a cadenceCOST 3.2Treat idle resources as defects rather than as slackCOST 3.3Know which resizings are cheap and which are migrationsCOST 3.4Right-size the shape, not only the size
COST 3.1 Compare provisioned capacity against measured usage on a cadence
Section titled “COST 3.1 Compare provisioned capacity against measured usage on a cadence”Risk if not established: Medium
Without a cadence, sizing is revisited when a budget review forces it, which is the most expensive moment to do it and the one where the decisions are worst.
Compare the two numbers per component: what is provisioned and what is used at peak. The gap is the opportunity, and it is usually larger than expected because sizing decisions are made once with a safety margin and then inherited.
Measure the peak rather than the average, at a resolution that catches bursts. PERF 3.1 makes
the same point for the opposite reason: an average that looks comfortable can hide a saturation,
and a peak that looks alarming can be a five-second burst that nothing depends on.
Do it alongside the performance review under PERF 9.1. The two look at the same measurements and
reach conclusions that need reconciling, and reconciling them in one conversation is cheaper than
discovering the conflict afterwards.
On STACKIT. The Cost
Dashboard supplies the
spend side per project, with monthly, quarterly, half-yearly, yearly and user-defined ranges. The
usage side comes from
Observability and
your instrumentation under OPS 7. Neither is useful alone: spend without utilization tells you
what you pay and not whether it is warranted.
Tradeoffs. Operational Excellence. A recurring review costs time from people who could be building, and its output is frequently a change that carries its own risk.
Verify. For your five largest cost items, what is provisioned and what is used at peak? When was that comparison last made?
COST 3.2 Treat idle resources as defects rather than as slack
Section titled “COST 3.2 Treat idle resources as defects rather than as slack”Risk if not established: Medium
There is a difference between headroom and waste, and it is whether anyone decided. Headroom is
capacity reserved for a stated reason under REL 7.1. Waste is capacity nobody is aware of.
The recurring finds: volumes detached from any instance, instances stopped but still billing, load balancers with no backends, snapshots from a migration that finished, environments from a project that ended, and IP addresses reserved and unused.
Each is small and they accumulate, and none of them will ever be found by looking at the largest
line items, which is why COST 7.2 argues for looking at the spend that buys least rather than
the spend that is biggest.
Make finding them recurring rather than heroic. A quarterly sweep that deletes twenty forgotten resources is worth more than an annual project, because the resources were created continuously.
On STACKIT. One billing behaviour is the most common source of an instance that costs money
while doing nothing: a machine that is merely stopped
keeps its resource reservation and continues to be billed for it, while a shelved machine does
not. Both look like the instance is off. COST 5.2 carries the distinction in full, with the
service-certificate wording it rests on.
Tradeoffs. Reliability. Deleting something that turns out to be needed is the risk, which
is why the ownership from COST 2.4 comes first. An unclaimed resource is easier to delete
confidently than an unlabelled one.
Verify. How many volumes in your estate are not attached to anything? How many instances are stopped rather than shelved?
COST 3.3 Know which resizings are cheap and which are migrations
Section titled “COST 3.3 Know which resizings are cheap and which are migrations”Risk if not established: Medium
Right-sizing assumes the size can be changed. Where it cannot, the decision is not whether to resize but whether to migrate, and that changes the arithmetic entirely.
An over-provisioned component whose size is a setting should be corrected as soon as it is noticed. One whose correction requires a migration needs the saving weighed against the effort and the risk, and the answer is sometimes to leave it.
This is PERF 3.2 from the cost side and the sort is the same: a setting, a replacement, or a
migration. Making it once and recording it serves both pillars.
The asymmetry is worth naming. Sizing up under pressure and sizing down at leisure have very different urgency, so an irreversible axis should be sized with more margin than a reversible one, which means accepting some waste deliberately.
On STACKIT. Compute Engine offers machine types as fixed variants, and a configuration cannot be adapted beyond them. Changing type is a replacement of the instance, which is disruptive and bounded.
Managed database performance classes are the case where a size correction is a migration: a class that is too small requires cloning to a new instance. Downsizing an over-provisioned class carries the same cost, which means an over-provisioned database is frequently cheaper to leave than to correct.
Tradeoffs. Cost Optimization, against itself. Accepting known waste on an irreversible axis is a deliberate cost, justified by the risk and effort of the alternative.
Verify. For each over-provisioned component, is correcting it a setting, a replacement or a migration? For the migrations, is the saving worth the work?
COST 3.4 Right-size the shape, not only the size
Section titled “COST 3.4 Right-size the shape, not only the size”Risk if not established: Medium
Sizing on one dimension means over-provisioning every other dimension to obtain enough of the one that binds. A memory-bound service on a CPU-weighted instance pays for cores it does not use.
Establish which resource actually limits the component, which is the measurement PERF 3.3 and
PERF 7.1 both need, then choose an option weighted toward it. The saving comes from no longer
buying the dimensions that were never the constraint.
Storage tiers are the dimension most often ignored, because they are not visible in the vCPU and RAM figures that dominate a sizing conversation. A workload on a high I/O tier that performs few operations is paying for throughput it does not consume.
Watch for components whose shape changed. A service that was CPU-bound before a cache was added may be memory-bound after it, and the instance chosen for the old shape is now wrong in a new direction.
On STACKIT. The machine type families are organized on exactly this axis, by vCPU-to-RAM ratio from CPU-weighted through general purpose to memory-weighted. Choosing the family is choosing the shape and matters more than choosing the size within one.
Managed databases separate the two dimensions: the flavor covers compute and memory, the performance class covers I/O, and they are chosen independently. That is useful and means both can be over-provisioned independently, so both belong in the review.
Tradeoffs. None material. Matching the shape reduces cost and improves performance at once, which is one of the few places where the two pillars agree without qualification.
Verify. For your largest component, which resource is the binding constraint? Does its shape weight that resource, or was it chosen on total size?
Related
Section titled “Related”PERF 3Selection and sizing, the same decisions from the performance sidePERF 9.1Lifecycle review, which should happen in the same conversationCOST 5Environments, where the largest idle capacity usually sitsREL 7.1Headroom, which is the deliberate version of unused capacitySUS 2Right-sizing, the same lever with a different stopping point