Zum Inhalt springen
Beta

COST 8. How do you make cost visible to the people who cause it?

Zuletzt aktualisiert am

Cost data that reaches only finance produces reports. Cost data that reaches the engineers whose decisions create it produces different decisions, and that is the whole mechanism by which this pillar works.

The gap is structural. The person who could delete the forgotten environment does not receive a bill, and the person who receives the bill cannot tell which environment is forgotten. Closing it requires the attribution from COST 2 and a delivery path that reaches people who are not looking for it.

  • COST 8.1 Put cost in front of the engineers whose decisions produce it
  • COST 8.2 Alert on anomalies rather than waiting for an invoice
  • COST 8.3 Show cost as a trend and per unit, not as a total
  • COST 8.4 Know the latency of your cost data and design around it

COST 8.1 Put cost in front of the engineers whose decisions produce it

Section titled “COST 8.1 Put cost in front of the engineers whose decisions produce it”

Risk if not established: Medium

An aggregate figure in a monthly finance review changes nothing, because nobody in that meeting can act on it and nobody who can act on it is in the meeting.

Deliver it where the work happens: a dashboard the team already looks at, a figure in the channel they use, a number attached to the project they own. The property that matters is that it arrives without anyone going to find it.

Attach it to something the team recognizes as theirs. A cost for a product they own is actionable; a share of an organizational total is not, which is the practical reason COST 2.1 insists the hierarchy reflect ownership.

Present it without blame. Cost visibility that arrives as criticism produces defensiveness and creative accounting rather than smaller bills, in the same way OPS 1.2 describes for incidents.

On STACKIT. The Cost Dashboard breaks costs down per project, which is the right granularity for a team that owns projects. Access to it follows the access model, so a team with access to its own projects can see its own costs without seeing everything.

Where the delivery has to be automated rather than pulled, the Cost API retrieves data per project and per customer account, which is what turns a dashboard somebody could look at into a figure that arrives.

Tradeoffs. Operational Excellence. Building the delivery is real work, and a cost report nobody reads is worse than none because it creates the impression that visibility exists.

Verify. Can each team see what it spends without asking anyone? When did a team last change something because of a cost figure they saw?


COST 8.2 Alert on anomalies rather than waiting for an invoice

Section titled “COST 8.2 Alert on anomalies rather than waiting for an invoice”

Risk if not established: High

A monthly invoice detects a runaway cost up to a month after it started. The recurring case is a misconfiguration that provisions continuously, a job in a retry loop, or a test environment created for an afternoon and left running.

Alert on the change rather than the level. A component whose cost doubles is worth knowing about regardless of whether the absolute figure is large, because the doubling is the signal that something changed unintentionally.

Set expectations per project so the alert has something to compare against. That is the same information COST 1 produced, which is one of the returns on having built a model at all.

Route it to the owner rather than to finance. An anomaly alert reaching someone who cannot act on it has converted a fast signal into a slow one.

On STACKIT. One property of the platform’s cost data shapes what is achievable here, and it is worth designing around rather than discovering. The Cost Dashboard provides no cost data for the current day, and the previous day’s data becomes available after 07:30 UTC.

So the fastest possible detection is next-day rather than same-hour. A runaway provisioned at nine in the morning is visible the following morning at the earliest, and a weekend mistake is visible on Monday. That is considerably better than a monthly invoice and it is not real time, which means guardrails matter more than detection.

The guardrails available are quotas and the bounds on autoscaling from PERF 5.2. A maximum that prevents a runaway from provisioning indefinitely is worth more than an alert that arrives a day later, precisely because the alert cannot arrive sooner.

Tradeoffs. Operational Excellence. Anomaly thresholds need tuning, and a noisy cost alert gets muted like any other, which is the fatigue problem from REL 10.2.

Verify. If a misconfiguration started provisioning resources this afternoon, when would somebody find out? What would have limited the damage in the meantime?


COST 8.3 Show cost as a trend and per unit, not as a total

Section titled “COST 8.3 Show cost as a trend and per unit, not as a total”

Risk if not established: Medium

A total answers whether spending went up. It does not answer whether that was justified, and a growing business with growing costs is not a problem.

Two presentations make the figure interpretable. The trend, because direction and rate matter more than the current value, and because a slow rise is invisible in a monthly comparison and obvious in a yearly one. The unit cost, meaning cost per customer, per transaction or per whatever the business counts, because that separates growth from inefficiency.

Unit cost is the one that changes conversations. A total that rose twenty percent while unit cost fell ten percent is a good month, and no view of the total alone can say so.

It also detects the opposite: a total that is flat while unit cost rises means the business is shrinking or the system is getting less efficient, and both are worth knowing early.

On STACKIT. The Cost Dashboard offers monthly, quarterly, half-yearly, yearly and user-defined ranges, which covers the trend view directly. The longer ranges are the ones that reveal slow growth, since a month-on-month view of a gradually rising line looks flat.

The unit figure is not on the platform, because the platform does not know what your business counts. Combining cost per project from the Cost API with a business metric from your own instrumentation under OPS 7.3 is what produces it.

Tradeoffs. Operational Excellence. Unit cost requires a business metric to be available and trustworthy, which is instrumentation work and an agreement about what to count.

Verify. What is your cost per unit of business value, and has it risen or fallen over the last year? If you cannot answer, which of the two inputs is missing?


COST 8.4 Know the latency of your cost data and design around it

Section titled “COST 8.4 Know the latency of your cost data and design around it”

Risk if not established: Medium

Cost data is never real time, and treating it as though it were produces a control that responds after the event it was meant to catch.

Establish the actual latency and design the response to it. Where data arrives daily, a daily anomaly check is the fastest useful control and anything more frequent is noise. Where the response has to be faster than the data, the answer is a preventive limit rather than a detective one.

That distinction is the practical output of this best practice. Detection tells you what happened; a quota tells you how bad it can get. Where detection is slow, the quota is doing most of the work and deserves the attention.

Set limits deliberately rather than accepting defaults. A quota that exists to prevent accidental overprovisioning is a cost control, and one set generously to avoid inconvenience is not.

On STACKIT. The Cost Dashboard states its own latency: no data for the current day, previous day available after 07:30 UTC. That figure is the input to every decision in this best practice, and having it stated saves you inferring it from when the numbers stop moving.

The preventive side comes from project quotas , which cover IaaS and Cloud Foundry resources rather than every service, from the autoscaling bounds in PERF 5.2, and from the documented service limits such as those for Kubernetes Engine . Each of those bounds what a mistake can cost before anybody sees it.

Tradeoffs. Operational Excellence. Tight quotas produce requests to raise them, which is friction. That friction is the control working, and it is the same trade SEC 5.1 makes for permissions.

Verify. How stale is your cost data when you look at it? What is the maximum a single misconfiguration could cost before the first cost signal arrives?


  • COST 2 Attribution, which decides whether a cost can be delivered to an owner
  • COST 1 Cost model, which supplies the expectation an anomaly is measured against
  • COST 9 Review cadence, the slower counterpart to anomaly alerting
  • PERF 5.4 Scaling cost, which is a cost control as well as a stability one
  • REL 10.2 Alerting, whose fatigue problem applies here identically