COST 8. How do you make cost visible to the people who cause it?
Zuletzt aktualisiert am
Cost data that reaches only finance produces reports. Cost data that reaches the engineers whose decisions create it produces different decisions, and that is the whole mechanism by which this pillar works.
The gap is structural. The person who could delete the forgotten environment does not receive a
bill, and the person who receives the bill cannot tell which environment is forgotten. Closing it
requires the attribution from COST 2 and a delivery path that reaches people who are not looking
for it.
Best practices
Section titled “Best practices”COST 8.1Put cost in front of the engineers whose decisions produce itCOST 8.2Alert on anomalies rather than waiting for an invoiceCOST 8.3Show cost as a trend and per unit, not as a totalCOST 8.4Know the latency of your cost data and design around it
COST 8.1 Put cost in front of the engineers whose decisions produce it
Section titled “COST 8.1 Put cost in front of the engineers whose decisions produce it”Risk if not established: Medium
An aggregate figure in a monthly finance review changes nothing, because nobody in that meeting can act on it and nobody who can act on it is in the meeting.
Deliver it where the work happens: a dashboard the team already looks at, a figure in the channel they use, a number attached to the project they own. The property that matters is that it arrives without anyone going to find it.
Attach it to something the team recognizes as theirs. A cost for a product they own is actionable;
a share of an organizational total is not, which is the practical reason COST 2.1 insists the
hierarchy reflect ownership.
Present it without blame. Cost visibility that arrives as criticism produces defensiveness and
creative accounting rather than smaller bills, in the same way OPS 1.2 describes for incidents.
On STACKIT. The Cost Dashboard breaks costs down per project, which is the right granularity for a team that owns projects. Access to it follows the access model, so a team with access to its own projects can see its own costs without seeing everything.
Where the delivery has to be automated rather than pulled, the Cost API retrieves data per project and per customer account, which is what turns a dashboard somebody could look at into a figure that arrives.
Tradeoffs. Operational Excellence. Building the delivery is real work, and a cost report nobody reads is worse than none because it creates the impression that visibility exists.
Verify. Can each team see what it spends without asking anyone? When did a team last change something because of a cost figure they saw?
COST 8.2 Alert on anomalies rather than waiting for an invoice
Section titled “COST 8.2 Alert on anomalies rather than waiting for an invoice”Risk if not established: High
A monthly invoice detects a runaway cost up to a month after it started. The recurring case is a misconfiguration that provisions continuously, a job in a retry loop, or a test environment created for an afternoon and left running.
Alert on the change rather than the level. A component whose cost doubles is worth knowing about regardless of whether the absolute figure is large, because the doubling is the signal that something changed unintentionally.
Set expectations per project so the alert has something to compare against. That is the same
information COST 1 produced, which is one of the returns on having built a model at all.
Route it to the owner rather than to finance. An anomaly alert reaching someone who cannot act on it has converted a fast signal into a slow one.
On STACKIT. One property of the platform’s cost data shapes what is achievable here, and it is worth designing around rather than discovering. The Cost Dashboard provides no cost data for the current day, and the previous day’s data becomes available after 07:30 UTC.
So the fastest possible detection is next-day rather than same-hour. A runaway provisioned at nine in the morning is visible the following morning at the earliest, and a weekend mistake is visible on Monday. That is considerably better than a monthly invoice and it is not real time, which means guardrails matter more than detection.
The guardrails available are quotas and the bounds on autoscaling from PERF 5.2. A maximum that
prevents a runaway from provisioning indefinitely is worth more than an alert that arrives a day
later, precisely because the alert cannot arrive sooner.
Tradeoffs. Operational Excellence. Anomaly thresholds need tuning, and a noisy cost alert
gets muted like any other, which is the fatigue problem from REL 10.2.
Verify. If a misconfiguration started provisioning resources this afternoon, when would somebody find out? What would have limited the damage in the meantime?
COST 8.3 Show cost as a trend and per unit, not as a total
Section titled “COST 8.3 Show cost as a trend and per unit, not as a total”Risk if not established: Medium
A total answers whether spending went up. It does not answer whether that was justified, and a growing business with growing costs is not a problem.
Two presentations make the figure interpretable. The trend, because direction and rate matter more than the current value, and because a slow rise is invisible in a monthly comparison and obvious in a yearly one. The unit cost, meaning cost per customer, per transaction or per whatever the business counts, because that separates growth from inefficiency.
Unit cost is the one that changes conversations. A total that rose twenty percent while unit cost fell ten percent is a good month, and no view of the total alone can say so.
It also detects the opposite: a total that is flat while unit cost rises means the business is shrinking or the system is getting less efficient, and both are worth knowing early.
On STACKIT. The Cost Dashboard offers monthly, quarterly, half-yearly, yearly and user-defined ranges, which covers the trend view directly. The longer ranges are the ones that reveal slow growth, since a month-on-month view of a gradually rising line looks flat.
The unit figure is not on the platform, because the platform does not know what your business
counts. Combining cost per project from the Cost
API with a
business metric from your own instrumentation under OPS 7.3 is what produces it.
Tradeoffs. Operational Excellence. Unit cost requires a business metric to be available and trustworthy, which is instrumentation work and an agreement about what to count.
Verify. What is your cost per unit of business value, and has it risen or fallen over the last year? If you cannot answer, which of the two inputs is missing?
COST 8.4 Know the latency of your cost data and design around it
Section titled “COST 8.4 Know the latency of your cost data and design around it”Risk if not established: Medium
Cost data is never real time, and treating it as though it were produces a control that responds after the event it was meant to catch.
Establish the actual latency and design the response to it. Where data arrives daily, a daily anomaly check is the fastest useful control and anything more frequent is noise. Where the response has to be faster than the data, the answer is a preventive limit rather than a detective one.
That distinction is the practical output of this best practice. Detection tells you what happened; a quota tells you how bad it can get. Where detection is slow, the quota is doing most of the work and deserves the attention.
Set limits deliberately rather than accepting defaults. A quota that exists to prevent accidental overprovisioning is a cost control, and one set generously to avoid inconvenience is not.
On STACKIT. The Cost Dashboard states its own latency: no data for the current day, previous day available after 07:30 UTC. That figure is the input to every decision in this best practice, and having it stated saves you inferring it from when the numbers stop moving.
The preventive side comes from project
quotas , which cover IaaS and Cloud Foundry resources rather than every service, from the autoscaling bounds in
PERF 5.2, and from
the documented service limits such as those for Kubernetes
Engine .
Each of those bounds what a mistake can cost before anybody sees it.
Tradeoffs. Operational Excellence. Tight quotas produce requests to raise them, which is
friction. That friction is the control working, and it is the same trade SEC 5.1 makes for
permissions.
Verify. How stale is your cost data when you look at it? What is the maximum a single misconfiguration could cost before the first cost signal arrives?
Related
Section titled “Related”COST 2Attribution, which decides whether a cost can be delivered to an ownerCOST 1Cost model, which supplies the expectation an anomaly is measured againstCOST 9Review cadence, the slower counterpart to anomaly alertingPERF 5.4Scaling cost, which is a cost control as well as a stability oneREL 10.2Alerting, whose fatigue problem applies here identically