SUS 1. How do you measure the resource consumption of your workload?
Zuletzt aktualisiert am
Every pillar begins with measurement. This one begins with whether measurement is possible, which is a different and less comfortable question.
Cost is on an invoice. Latency is on a timer. Environmental footprint depends on facts that are not visible from inside a virtual machine: how the hardware is shared, how the facility is powered, and what the marginal energy source was at the moment your job ran. What you can observe are proxies, and the honest starting point is knowing which ones you have and what each of them omits.
Best practices
Section titled “Best practices”SUS 1.1Establish what you can actually measure before setting a target against itSUS 1.2Use proxies you can defend, and state what each one omitsSUS 1.3Separate what the provider controls from what you controlSUS 1.4Do not optimize a metric that does not track physical consumption
SUS 1.1 Establish what you can actually measure before setting a target against it
Section titled “SUS 1.1 Establish what you can actually measure before setting a target against it”Risk if not established: Medium
A target set against a metric nobody can produce is a reporting exercise waiting to happen. Start from the signals that exist rather than from the figure you would like to report.
The order matters. Establish the available measurements, decide which of them correlate with consumption, then set targets against those. Reversing it produces a target that gets satisfied by changing how it is calculated.
Distinguish three kinds of figure, because they have very different reliability. Directly measured quantities such as allocated capacity, storage volume and running hours. Derived quantities such as an emissions estimate produced from those by applying factors. And provider figures such as facility efficiency, which describe the infrastructure rather than your workload.
Only the first is yours to verify. The other two depend on methodology you did not choose and may not be able to inspect, which is worth knowing before a number ends up in a report with your name on it.
On STACKIT. Two provider-level facts are published and both are relevant, because they set the context your workload sits in. All data centres are operated on green electricity from regional sources in Germany and Austria, and the facility in Ostermiething reports a PUE of 1.1 , which is a strong figure by European standards.
The same page states that integrated tools let companies measure the energy consumption and emissions of their cloud services. That capability would change this question substantially, and it is the single most consequential thing to confirm before designing a measurement approach.
Establish four things about that reporting before an obligation rests on it: where it is
available, what it reports, at what granularity, and by what methodology. A figure whose
methodology you cannot state is not evidence for a disclosure. Until those four are answered, the
proxies in SUS 1.2 are what a measurement approach should stand on.
Tradeoffs. Operational Excellence. Time spent establishing what is measurable before measuring anything, which feels like delay and prevents building on a figure that turns out not to exist.
Verify. Which consumption figures can you produce for your workload today? For each, where does it come from and who calculated it?
SUS 1.2 Use proxies you can defend, and state what each one omits
Section titled “SUS 1.2 Use proxies you can defend, and state what each one omits”Risk if not established: Medium
Where direct measurement is unavailable, proxies are the working answer. They are legitimate provided their limits are stated, and misleading when presented as measurements.
The proxies that correlate reasonably with consumption and that you can produce yourself:
- Allocated capacity, meaning vCPU, memory and storage provisioned rather than used. It correlates with hardware occupied, which is the physical quantity.
- The gap between allocated and used, which is the waste this pillar addresses and the figure most likely to produce action.
- Running hours, particularly for resources that could be off, which is
SUS 6. - Stored volume and its growth rate, which correlates with hardware occupied over time.
Each omits something. Allocated capacity says nothing about the efficiency of the hardware or the facility. Running hours treat an idle instance and a saturated one identically. None of them capture embodied footprint, which was paid at manufacture regardless of use.
State the omissions where the figure is used. A proxy with its limits attached is an honest input to a decision; the same proxy presented as a measurement is a claim you cannot support.
On STACKIT. These proxies come from the same sources the other pillars already use: utilization from Observability , storage volume and growth from the services holding it, and allocated capacity from the resource inventory.
Cost is the most available proxy and the most misleading one, which is why it has its own best
practice in SUS 1.4. The
Cost Dashboard makes it
easy to reach for, and its per-project breakdown is genuinely useful for locating waste even where
it is a poor proxy for consumption.
Tradeoffs. Operational Excellence. Maintaining proxy measurements is work with no external reporting value until the direct figures exist, which makes it easy to defer.
Verify. For each consumption figure you report, is it measured, derived or a proxy? Where is the limitation of each stated?
SUS 1.3 Separate what the provider controls from what you control
Section titled “SUS 1.3 Separate what the provider controls from what you control”Risk if not established: Medium
Sustainability figures mix two very different things: how efficiently the infrastructure runs, and how much of it you use. Conflating them produces a number nobody can act on, because half of it is not yours.
The provider layer covers energy sourcing, facility efficiency, cooling, and hardware lifecycle. Those are procurement decisions rather than architecture decisions, and they belong in the choice of provider rather than in a design review.
Your layer is the amount of infrastructure you occupy and for how long. That is what every other question in this pillar addresses, and it is the only part a workload design can change.
Keeping them separate has a practical benefit beyond clarity: it prevents the argument that efficiency at the provider layer removes the need for efficiency at yours. Green electricity makes consumption cleaner and does not make it free, and it certainly does not make idle capacity justified.
On STACKIT. The provider layer is documented and strong: green electricity from regional sources across all data centres, and a PUE of 1.1 in Ostermiething achieved with river water cooling and heat recovery.
That is worth knowing precisely because it tells you where your remaining lever is. When the
infrastructure layer is already efficient, the variable that is left is how much of it you occupy,
which is why this pillar’s centre of gravity is SUS 4 rather than anything the platform
provides.
Tradeoffs. None. This is a framing decision that costs nothing and prevents a category of unproductive argument.
Verify. In your sustainability reporting, which figures describe the provider’s infrastructure and which describe your workload? Can a reader tell them apart?
SUS 1.4 Do not optimize a metric that does not track physical consumption
Section titled “SUS 1.4 Do not optimize a metric that does not track physical consumption”Risk if not established: Medium
This is the failure mode the whole question exists to prevent. A workload optimized against a number that does not correspond to physical reality has achieved a reporting outcome, and distinguishing that from an environmental one is most of the discipline in this pillar.
Cost is the metric this happens with most often, because it is available, precise and looks authoritative. It diverges from consumption in specific and predictable ways, which the Sustainability tradeoffs set out: committed capacity is cheaper and equally allocated, cheaper storage tiers remove the pressure to delete without freeing any hardware, and cost optimization stops when spend is acceptable while consumption keeps mattering.
Test any metric before optimizing against it by asking what would happen if you improved it without changing anything physical. If that is possible, the metric is measuring an accounting artefact.
Where cost and consumption agree, use cost. It is easier to obtain, easier to fund and reaches decision-makers. Where they diverge, this pillar needs its own argument, and having named the divergence in advance is what makes that argument possible.
On STACKIT. The clearest divergence available on the platform is the shelving distinction from
COST 5.2, and it runs in the direction that favours sustainability. A machine that is stopped
without shelving keeps its resource reservation, which means it is still occupying capacity as
well as still being billed. Shelving releases both.
That makes it one of the few places where the cost signal and the physical signal point the same way, which is worth using rather than arguing about.
Tradeoffs. Cost Optimization, occasionally. Where the two diverge, following consumption rather than cost means spending money for an outcome that does not appear on the invoice.
Verify. Take your main sustainability metric. Could it improve without any change to how much hardware you occupy or for how long? If so, what is it actually measuring?
Related
Section titled “Related”SUS 2Right-sizing, the first place these measurements are appliedSUS 4Utilization density, the lever the provider layer leaves to youCOST 8Cost visibility, whose data is the most available and most misleading proxyOPS 7Observability, which supplies the utilization side- Sustainability tradeoffs, which sets out where cost and consumption part