SUS 2. How do you provision for real demand rather than for comfort?
Last updated on
COST 3 asks the same question about money and stops when the spend is acceptable. This question
continues past that point, because allocated capacity occupies hardware whether or not its price
has become tolerable.
The behaviour it addresses is not carelessness. Rounding up is rational when demand is uncertain, when the next size costs a little more and removes a risk, and when being short is visible while being generous is not. That rationality is exactly why the correction has to be deliberate.
Best practices
Section titled “Best practices”SUS 2.1Provision against measured usage rather than against the size that stops you thinkingSUS 2.2Treat rounding up as a decision with a consequenceSUS 2.3Distinguish headroom from slackSUS 2.4Continue past the point where cost stops caring
SUS 2.1 Provision against measured usage rather than against the size that stops you thinking
Section titled “SUS 2.1 Provision against measured usage rather than against the size that stops you thinking”Risk if not established: Medium
The comfortable size is the one large enough that nobody has to think about capacity again. It is chosen once, it works, and nothing about a working system draws attention to the gap between what is allocated and what is used.
Measure that gap per component. It is the single most useful figure in this pillar because it is directly actionable, and it is usually larger than expected for the reason above.
Use peak rather than average, at a resolution that catches bursts, which is the same measurement
PERF 3.1 and COST 3.1 need. Doing it once and sharing the result across three reviews is the
practical way to make this affordable.
Include the failure case in what counts as needed. Capacity that absorbs a zone loss under
REL 7.1 is required rather than spare, and treating it as waste is how a sustainability review
degrades reliability.
On STACKIT. Some of the gap is structural rather than a choice, and it is worth naming so that the review does not chase it.
Machine types
come as fixed variants, and a configuration cannot be adapted beyond them. Where measured demand falls between two variants, you take the larger one, so
allocation exceeds requirement by construction. That portion of the gap cannot be closed by better
sizing; it can only be closed by changing the shape so a smaller variant fits, which is SUS 2.2.
Tradeoffs. Reliability. Sizing closer to measured demand narrows the margin for demand
being underestimated, which is why REL 7.1 sets the headroom deliberately rather than leaving it
as a residue.
Verify. For your five largest components, what is allocated and what is used at peak? How much of that gap is the next variant up rather than a choice?
SUS 2.2 Treat rounding up as a decision with a consequence
Section titled “SUS 2.2 Treat rounding up as a decision with a consequence”Risk if not established: Medium
Every rounding decision is individually small and defensible. The sum of them across an estate is the largest single source of allocated-but-unused capacity, and no single decision will ever look like the problem.
Make the rounding visible rather than automatic. When measured demand sits between two options, that is a choice between accepting the larger allocation and doing something so the smaller one fits. The second is available more often than it is considered.
The ways the smaller option becomes viable: reducing the work under SUS 7, changing the resource
shape so the binding dimension is what you buy under PERF 3.3, splitting the workload so it fits
two small allocations rather than one large, or accepting a lower target under PERF 1.
Where you round up, note by how much. An estate where every component is one variant larger than it needs is carrying a consistent overhead that is invisible per component and substantial in total.
On STACKIT. The variant granularity determines how much rounding costs. Where the steps
between variants are large, the penalty for falling just above a threshold is large too, which
makes the shape decision from PERF 3.3 more valuable than it looks: choosing a family whose
ratio matches the workload means the binding dimension determines the variant rather than an
unrelated one.
Managed services have the same structure through flavors and performance classes, with the
additional constraint from PERF 3.2 that a database performance class cannot be changed in
place. That makes over-allocation on that axis harder to correct, which is an argument for getting
the shape right rather than for rounding up further.
Tradeoffs. Operational Excellence. Making the smaller option fit is engineering work, and accepting the larger allocation is free today. That asymmetry is why this needs to be a recorded decision rather than a default.
Verify. For components where you chose the larger variant, what would have been needed to make the smaller one fit? Was that considered?
SUS 2.3 Distinguish headroom from slack
Section titled “SUS 2.3 Distinguish headroom from slack”Risk if not established: Medium
Both look identical in a utilization graph. Headroom is capacity reserved for a stated reason:
absorbing a peak under REL 7.1, surviving a zone loss under REL 4.1, or covering the time
scaling takes under PERF 5.4. Slack is capacity nobody decided on.
Only slack is a target for this question, and treating them as the same thing produces a sustainability review that reduces reliability, which is the tension the Sustainability tradeoffs describe.
Sort them explicitly. For each component with unused capacity, the question is whether a stated requirement accounts for it. Where one does, the capacity is justified and the review moves on. Where none does, it is slack regardless of how comfortable it feels.
The uncomfortable finding is usually that most of the unused capacity has no stated requirement
behind it, because the targets in REL 1 and PERF 1 were never set. Where that is the case,
this question cannot be answered until those are.
On STACKIT. Headroom sizing depends on how quickly capacity can be added, which is why REL 7.4 and PERF 5.4 both ask for the provisioning time to be measured. A component that scales in
seconds needs less standing headroom than one that takes minutes, and the difference is allocated
hardware standing idle.
That measurement is workload-specific and worth recording once, because it is an input to reliability, performance and this pillar simultaneously.
Tradeoffs. Reliability, directly. Every reduction in headroom is a reduction in the margin for something unexpected, which is why the sort is between slack and headroom rather than a uniform cut.
Verify. For each component, how much unused capacity is there and which stated requirement justifies it? What proportion has no requirement behind it?
SUS 2.4 Continue past the point where cost stops caring
Section titled “SUS 2.4 Continue past the point where cost stops caring”Risk if not established: Medium
This best practice is the reason the pillar exists separately from COST 3. Cost optimization has
a stopping condition, which is that spend has become acceptable. Consumption does not.
Three specific cases where cost stops and consumption should not:
Committed or reserved capacity is cheaper per unit and equally allocated. The financial pressure to reduce it disappears and the hardware occupancy does not change at all.
Cheap resources attract less scrutiny in proportion to their price, so an oversized allocation of something inexpensive survives every cost review while occupying the same physical capacity as an expensive one of the same size.
An acceptable total ends a cost review. A workload whose spend is within budget is not examined further, regardless of how much of what it pays for is unused.
Where cost and consumption agree, act on the cost argument. It is easier to fund and the outcome
is identical, which is the practical advice in SUS 1.4. Reserve the separate argument for these
divergences, because that is where the consumption argument is the only one available.
On STACKIT. The
Cost Dashboard shows the
spend and not the occupancy, so a component whose price fell without its allocation changing looks
improved. Pairing it with utilization from
Observability is what
distinguishes the two, and it is the same pairing COST 3.1 needs.
Tradeoffs. Cost Optimization, in the sense that this best practice asks for work with no financial return, which makes it the part that has to be argued on consumption alone.
Verify. Name a component whose cost is acceptable and whose allocation exceeds its usage substantially. What would prompt anyone to change it?
Related
Section titled “Related”COST 3Right-sizing, the same lever with an earlier stopping pointPERF 3Selection and sizing, which decides the shape and the variantREL 7.1Headroom, which is the justified portion of unused capacitySUS 4Utilization density, the structural version of the same problem- Sustainability tradeoffs