---
id: COST08
pillar: cost-optimization
title: COST 8. How do you make cost visible to the people who cause it?
description: The engineer who could turn off the idle cluster never sees the invoice, and the person who sees the invoice does not know which cluster is the idle one.
status: draft
services: []
sidebar:
  order: 17
  label: Cost visibility
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/cost-08-cost-visibility/"
source_file: "docs/architecture/pillars/cost-optimization/cost-08-cost-visibility.mdx"
---

Cost data that reaches only finance produces reports. Cost data that reaches the engineers whose
decisions create it produces different decisions, and that is the whole mechanism by which this
pillar works.

The gap is structural. The person who could delete the forgotten environment does not receive a
bill, and the person who receives the bill cannot tell which environment is forgotten. Closing it
requires the attribution from [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) and a delivery path that reaches people who are not looking
for it.

## Best practices

- [`COST 8.1`](/architecture/pillars/cost-optimization/cost-08-cost-visibility/#cost-81-put-cost-in-front-of-the-engineers-whose-decisions-produce-it) Put cost in front of the engineers whose decisions produce it
- [`COST 8.2`](/architecture/pillars/cost-optimization/cost-08-cost-visibility/#cost-82-alert-on-anomalies-rather-than-waiting-for-an-invoice) Alert on anomalies rather than waiting for an invoice
- [`COST 8.3`](/architecture/pillars/cost-optimization/cost-08-cost-visibility/#cost-83-show-cost-as-a-trend-and-per-unit-not-as-a-total) Show cost as a trend and per unit, not as a total
- [`COST 8.4`](/architecture/pillars/cost-optimization/cost-08-cost-visibility/#cost-84-know-the-latency-of-your-cost-data-and-design-around-it) Know the latency of your cost data and design around it

---

## COST 8.1 Put cost in front of the engineers whose decisions produce it

**Risk if not established:** Medium

An aggregate figure in a monthly finance review changes nothing, because nobody in that meeting
can act on it and nobody who can act on it is in the meeting.

Deliver it where the work happens: a dashboard the team already looks at, a figure in the channel
they use, a number attached to the project they own. The property that matters is that it arrives
without anyone going to find it.

Attach it to something the team recognizes as theirs. A cost for a product they own is actionable;
a share of an organizational total is not, which is the practical reason [`COST 2.1`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-21-make-the-resource-hierarchy-reflect-who-pays) insists the
hierarchy reflect ownership.

Present it without blame. Cost visibility that arrives as criticism produces defensiveness and
creative accounting rather than smaller bills, in the same way [`OPS 1.2`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-12-make-incident-review-blameless-in-practice-not-only-in-policy) describes for incidents.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost
Dashboard</LinkChip> breaks costs down
per project, which is the right granularity for a team that owns projects. Access to it follows
the access model, so a team with access to its own projects can see its own costs without seeing
everything.

Where the delivery has to be automated rather than pulled, the <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/how-tos/retrieve-cost-data/">Cost
API</LinkChip> retrieves
data per project and per customer account, which is what turns a dashboard somebody could look at
into a figure that arrives.

**Tradeoffs.** **Operational Excellence.** Building the delivery is real work, and a cost report
nobody reads is worse than none because it creates the impression that visibility exists.

**Verify.** Can each team see what it spends without asking anyone? When did a team last change
something because of a cost figure they saw?

---

## COST 8.2 Alert on anomalies rather than waiting for an invoice

**Risk if not established:** High

A monthly invoice detects a runaway cost up to a month after it started. The recurring case is a
misconfiguration that provisions continuously, a job in a retry loop, or a test environment
created for an afternoon and left running.

Alert on the change rather than the level. A component whose cost doubles is worth knowing about
regardless of whether the absolute figure is large, because the doubling is the signal that
something changed unintentionally.

Set expectations per project so the alert has something to compare against. That is the same
information [`COST 1`](/architecture/pillars/cost-optimization/cost-01-cost-model/) produced, which is one of the returns on having built a model at all.

Route it to the owner rather than to finance. An anomaly alert reaching someone who cannot act on
it has converted a fast signal into a slow one.

**On STACKIT.** One property of the platform's cost data shapes what is achievable here, and it is
worth designing around rather than discovering. The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost
Dashboard</LinkChip> provides **no
cost data for the current day**, and the previous day's data becomes available after 07:30 UTC.

So the fastest possible detection is next-day rather than same-hour. A runaway provisioned at nine
in the morning is visible the following morning at the earliest, and a weekend mistake is visible
on Monday. That is considerably better than a monthly invoice and it is not real time, which means
guardrails matter more than detection.

The guardrails available are quotas and the bounds on autoscaling from [`PERF 5.2`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-52-scale-on-signals-that-predict-saturation-rather-than-confirm-it). A maximum that
prevents a runaway from provisioning indefinitely is worth more than an alert that arrives a day
later, precisely because the alert cannot arrive sooner.

**Tradeoffs.** **Operational Excellence.** Anomaly thresholds need tuning, and a noisy cost alert
gets muted like any other, which is the fatigue problem from [`REL 10.2`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-102-alert-on-user-visible-symptoms-rather-than-on-component-metrics-alone).

**Verify.** If a misconfiguration started provisioning resources this afternoon, when would
somebody find out? What would have limited the damage in the meantime?

---

## COST 8.3 Show cost as a trend and per unit, not as a total

**Risk if not established:** Medium

A total answers whether spending went up. It does not answer whether that was justified, and a
growing business with growing costs is not a problem.

Two presentations make the figure interpretable. **The trend**, because direction and rate matter
more than the current value, and because a slow rise is invisible in a monthly comparison and
obvious in a yearly one. **The unit cost**, meaning cost per customer, per transaction or per
whatever the business counts, because that separates growth from inefficiency.

Unit cost is the one that changes conversations. A total that rose twenty percent while unit cost
fell ten percent is a good month, and no view of the total alone can say so.

It also detects the opposite: a total that is flat while unit cost rises means the business is
shrinking or the system is getting less efficient, and both are worth knowing early.

**On STACKIT.** The
<LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost Dashboard</LinkChip> offers
monthly, quarterly, half-yearly, yearly and user-defined ranges, which covers the trend view
directly. The longer ranges are the ones that reveal slow growth, since a month-on-month view of a
gradually rising line looks flat.

The unit figure is not on the platform, because the platform does not know what your business
counts. Combining cost per project from the <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/how-tos/retrieve-cost-data/">Cost
API</LinkChip> with a
business metric from your own instrumentation under [`OPS 7.3`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-73-instrument-for-the-questions-you-cannot-predict) is what produces it.

**Tradeoffs.** **Operational Excellence.** Unit cost requires a business metric to be available
and trustworthy, which is instrumentation work and an agreement about what to count.

**Verify.** What is your cost per unit of business value, and has it risen or fallen over the last
year? If you cannot answer, which of the two inputs is missing?

---

## COST 8.4 Know the latency of your cost data and design around it

**Risk if not established:** Medium

Cost data is never real time, and treating it as though it were produces a control that responds
after the event it was meant to catch.

Establish the actual latency and design the response to it. Where data arrives daily, a daily
anomaly check is the fastest useful control and anything more frequent is noise. Where the
response has to be faster than the data, the answer is a preventive limit rather than a detective
one.

That distinction is the practical output of this best practice. Detection tells you what happened;
a quota tells you how bad it can get. Where detection is slow, the quota is doing most of the work
and deserves the attention.

Set limits deliberately rather than accepting defaults. A quota that exists to prevent accidental
overprovisioning is a cost control, and one set generously to avoid inconvenience is not.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost
Dashboard</LinkChip> states its own
latency: no data for the current day, previous day available after 07:30 UTC. That figure is the
input to every decision in this best practice, and having it stated saves you inferring it from
when the numbers stop moving.

The preventive side comes from <LinkChip href="https://docs.stackit.cloud/platform/resource-manager/basics/projects/">project
quotas</LinkChip>, which cover IaaS and Cloud Foundry resources rather than every service, from the autoscaling bounds in
[`PERF 5.2`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-52-scale-on-signals-that-predict-saturation-rather-than-confirm-it), and from
the documented service limits such as <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/operations/quotas-and-limits/">those for Kubernetes
Engine</LinkChip>.
Each of those bounds what a mistake can cost before anybody sees it.

**Tradeoffs.** **Operational Excellence.** Tight quotas produce requests to raise them, which is
friction. That friction is the control working, and it is the same trade [`SEC 5.1`](/architecture/pillars/security/sec-05-least-privilege/#sec-51-grant-the-narrowest-role-that-permits-the-work) makes for
permissions.

**Verify.** How stale is your cost data when you look at it? What is the maximum a single
misconfiguration could cost before the first cost signal arrives?

---

## Related

- [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) Attribution, which decides whether a cost can be delivered to an owner
- [`COST 1`](/architecture/pillars/cost-optimization/cost-01-cost-model/) Cost model, which supplies the expectation an anomaly is measured against
- [`COST 9`](/architecture/pillars/cost-optimization/cost-09-review-cadence/) Review cadence, the slower counterpart to anomaly alerting
- [`PERF 5.4`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-54-account-for-what-scaling-costs-in-time-and-in-money) Scaling cost, which is a cost control as well as a stability one
- [`REL 10.2`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-102-alert-on-user-visible-symptoms-rather-than-on-component-metrics-alone) Alerting, whose fatigue problem applies here identically
