COST 9. How do you review actual spend against the model on a fixed cadence?
Last updated on
Optimization done once produces a familiar pattern: a sharp improvement, eighteen months of steady erosion back to roughly where it started, and then another cost reduction project.
A cadence is what breaks that cycle. It does not have to be frequent. It has to exist, have an
owner, and compare reality against something that states what was expected, which is the model
from COST 1.
Best practices
Section titled “Best practices”COST 9.1Compare actual spend against the model rather than against last monthCOST 9.2Act on the delta by fixing the workload or by correcting the modelCOST 9.3Give the review an owner and a cadence proportional to the spendCOST 9.4Hold it alongside the performance and reliability reviews
COST 9.1 Compare actual spend against the model rather than against last month
Section titled “COST 9.1 Compare actual spend against the model rather than against last month”Risk if not established: Medium
Comparing this month with last month detects change and cannot detect a persistent error. A workload that has cost twice its estimate since launch looks stable in a month-on-month view, and stable is what nobody investigates.
Compare against the model. The delta is the finding, and its sign matters as much as its size: under budget usually means over-provisioned somewhere or a feature that never shipped, and both are worth knowing.
Diagnose the delta using the drivers from COST 1.3. A total that is forty percent high is a
mystery; a total that is forty percent high because data grew three times faster than modelled is
a decision about retention.
Where there is no model, the first review produces one. That is a legitimate outcome and better than skipping the review, since the second review then has something to compare against.
On STACKIT. The Cost
Dashboard supplies the
actual figures per project with monthly, quarterly, half-yearly, yearly and user-defined ranges.
The per-project breakdown is what makes the comparison possible at the granularity a model was
written at, which is the return on the hierarchy decision in COST 2.1.
For a review that runs on a schedule, the Cost API retrieves the same data programmatically, which turns preparation into a generated report rather than a manual assembly and makes the cadence survivable.
Tradeoffs. Little beyond the effort of maintaining a model, which COST 1 already argues for
on its own merits.
Verify. For your largest workload, what did the model predict and what did it actually cost last quarter? What explains the difference?
COST 9.2 Act on the delta by fixing the workload or by correcting the model
Section titled “COST 9.2 Act on the delta by fixing the workload or by correcting the model”Risk if not established: Medium
A review that produces observations and no changes is a reporting exercise. The output of this question is either a change to the workload or a change to the model, and doing neither is the failure mode.
Both outcomes are legitimate. If the model was wrong, correcting it makes the next comparison meaningful and is not an admission of anything. If the workload is wrong, the review has found what it was for.
Give each action an owner and a date, in the same way OPS 9.4 treats incident actions. Cost
findings compete with feature work and lose by default, and an action without an owner is an
observation with a deadline nobody holds.
Track what the review saved. It is one of the few places in this framework where the benefit is directly measurable, and it is what justifies the cadence continuing to exist when someone questions it.
On STACKIT. Whether an action can be taken depends on the reversibility sort from COST 3.3.
A review finding an over-provisioned setting produces an immediate change; one finding an
over-provisioned managed database performance class produces a decision about whether a migration
is worth it, since changing class requires cloning the instance.
Recording that distinction in the finding rather than in the discussion is what stops the same item appearing in three consecutive reviews.
Tradeoffs. Operational Excellence. Acting on findings means changes with their own risk, and the riskiest changes are frequently the ones with the largest savings.
Verify. From your last cost review, how many actions were identified, how many are complete, and how much did they save?
COST 9.3 Give the review an owner and a cadence proportional to the spend
Section titled “COST 9.3 Give the review an owner and a cadence proportional to the spend”Risk if not established: Medium
A review that belongs to everyone happens when someone has time, which under delivery pressure is never.
Name an owner who can convene it and who is accountable for it happening rather than for the
number going down. Those are different accountabilities, and confusing them produces a review that
optimizes for a reduction rather than for the right level of spend, which is what COST 7.3 warns
against.
Set the cadence from the spend and its volatility. A workload with a stable, modest cost warrants a quarterly look. One that is large, growing or newly launched warrants more. The cadence is a judgement rather than a standard.
Keep the review short by preparing it automatically. Most of the time in a badly run cost review goes into assembling the data, and none of that time produces a decision.
On STACKIT. The Cost API is what makes preparation automatic, retrieving per-project and per-customer-account data on a schedule. That turns the review into a discussion about a report rather than an exercise in producing one, which is the difference between a cadence that survives and one that lapses.
Generating the report is a good candidate for OPS 10.2, since it is recurring, mechanical and
currently manual in most organizations.
Tradeoffs. Cost Optimization, against itself in a small way: the review consumes time that is not free. A quarterly hour with a prepared report is a different proposition from a quarterly day spent building one.
Verify. Who owns your cost review, how often does it happen, and how much of the time goes into preparing the data rather than deciding anything?
COST 9.4 Hold it alongside the performance and reliability reviews
Section titled “COST 9.4 Hold it alongside the performance and reliability reviews”Risk if not established: Medium
The cost review, the sizing review under PERF 9.1 and the reliability targets look at the same
components and reach conclusions that conflict. Held separately, whichever runs first acts and the
others discover the consequence.
The specific failure is the one COST 7.4 describes: a cost review reduces something that a
reliability target depended on, nobody in the room knew about the target, and the workload’s
actual capability drifts from what the business assumed.
Holding them together costs one longer meeting and removes an entire class of that problem. It also surfaces the cases where the pillars agree, which are the easiest actions available: an idle environment is a cost finding, a sustainability finding and frequently a security finding at once.
Bring the flow ranking from REL 2.2. It is the input all three reviews use, and having it
present stops the conversation drifting into resources rather than value.
On STACKIT. The measurements come from two places that have to be looked at together: spend from the Cost Dashboard and utilization from Observability . Neither is sufficient alone, which is the practical reason a joint review is more than an organizational convenience.
Tradeoffs. Operational Excellence. More people in one meeting, harder to schedule, and a longer session. The alternative is a sequence of locally rational decisions that are collectively wrong.
Verify. When you last reduced a cost, who checked whether it affected a reliability or performance target? Was that person in the room, or consulted afterwards?
Related
Section titled “Related”COST 1Cost model, which supplies what the review compares againstCOST 7.4Reconciliation, the reason this review is held jointlyPERF 9.1Sizing review, which should happen in the same conversationCOST 3.3Reversibility, which decides whether a finding is actionableOPS 9.4Tracking actions, the same discipline applied to incident findings