---
id: COST07
pillar: cost-optimization
title: COST 7. How do you direct cost scrutiny by business value?
description: Optimizing a resource list from the top by price finds the largest line items, which is not at all the same as finding the spend that buys you the least.
status: draft
services: []
sidebar:
  order: 16
  label: Flow-based optimization
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/"
source_file: "docs/architecture/pillars/cost-optimization/cost-07-flow-based-optimization.mdx"
---

Cost reviews usually start by sorting resources by price and working down the list. That finds the
biggest numbers, and the biggest number is frequently the one doing the most valuable work.

The useful question is different. Not what costs most, but what returns least for what it costs.
Answering it requires knowing what each cost is for, which is why this question depends on the
attribution from [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) and the flow ranking from [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume).

## Best practices

- [`COST 7.1`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/#cost-71-rank-spend-by-the-flow-it-serves) Rank spend by the flow it serves
- [`COST 7.2`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/#cost-72-look-for-the-spend-that-buys-least-rather-than-the-largest-line-item) Look for the spend that buys least rather than the largest line item
- [`COST 7.3`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/#cost-73-accept-higher-cost-where-the-flow-justifies-it) Accept higher cost where the flow justifies it
- [`COST 7.4`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/#cost-74-reconcile-with-the-reliability-and-performance-decisions-rather-than-around-them) Reconcile with the reliability and performance decisions rather than around them

---

## COST 7.1 Rank spend by the flow it serves

**Risk if not established:** Medium

A cost sorted by resource type tells you that compute is expensive. A cost sorted by flow tells
you that the internal reporting path costs a third of your infrastructure, which is a decision
rather than an observation.

Use the ranking from [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume) and the flow-to-component map from [`REL 2.3`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-23-map-each-flow-to-every-component-and-dependency-it-touches). Those already exist
if the Reliability pillar has been worked through, and this is the third use of the same exercise
after reliability investment and performance targets.

Two things become visible that a resource view hides. Flows whose cost is disproportionate to
their importance, which are the targets. And costs that serve no flow at all, which are the
unclaimed resources from [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims).

Expect the mapping to be imperfect. Shared components serve several flows and dividing them
exactly is not worth the effort; an approximate split is enough to reveal a disproportion.

**On STACKIT.** Whether this is straightforward or laborious is decided by [`COST 2.1`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-21-make-the-resource-hierarchy-reflect-who-pays). Where the
project structure follows products and environments, the
<LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost Dashboard</LinkChip> already
groups costs close to the flow boundary. Where several products share a project, the split has to
be reconstructed from usage and the reconstruction is an estimate.

That is the practical consequence of the hierarchy decision being expensive to change later: it
determines whether this question takes an afternoon or a week.

**Tradeoffs.** **Operational Excellence.** Mapping cost to flows is analysis work that has to be
redone as the architecture changes, which is the same maintenance [`REL 2.4`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-24-keep-the-map-current-as-the-architecture-changes) describes for the flow
map itself.

**Verify.** What proportion of your infrastructure cost serves your highest-ranked flow? What
proportion serves flows nobody ranked?

---

## COST 7.2 Look for the spend that buys least rather than the largest line item

**Risk if not established:** Medium

The largest line item is usually the production database or the main compute tier, and it is
usually justified. Working down from the top produces a review that examines the well-justified
costs first and runs out of energy before reaching the questionable ones.

Invert it. The interesting costs are the ones with a poor ratio of value to price, and they
cluster in predictable places: non-production environments under [`COST 5`](/architecture/pillars/cost-optimization/cost-05-environments/), data nobody reads under
[`COST 6`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/), capacity provisioned for a peak that never arrived under [`COST 3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/), and components
serving flows that were deprioritized.

Ratio rather than size is the discipline. A small cost that buys nothing is a better target than a
large cost that buys a great deal, even though the second saves more if eliminated, because the
second cannot be eliminated.

Look for the costs that nobody would create today. An estate accumulates decisions that were
correct when made and would not be repeated, and those are the least contentious savings
available.

**On STACKIT.** The
<LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost Dashboard</LinkChip> gives the
price side per project. The value side is not on the platform: it comes from the flow ranking and
from knowing what each project is for, which is [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) again.

The specific case worth checking first is the difference between production and everything else.
In most estates the non-production total is larger than anyone expects, and it is visible
immediately where environments have their own projects.

**Tradeoffs.** Little. The main cost is that this ordering is less satisfying than working down
from the largest number, because the individual wins are smaller.

**Verify.** Sort your costs by how much value each buys rather than by price. What is at the
bottom of that list, and when was it last examined?

---

## COST 7.3 Accept higher cost where the flow justifies it

**Risk if not established:** Medium

A cost review that only ever reduces is not following value; it is following a target. Directing
spend by business value means spending more in some places, and a review that never concludes this
is not doing the job.

The cases where more is correct: a flow whose reliability target justifies redundancy under
[`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/), a flow whose performance target justifies headroom under [`PERF 3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/), and a component whose
managed equivalent costs more and removes work under [`COST 4.2`](/architecture/pillars/cost-optimization/cost-04-service-selection/#cost-42-compare-managed-against-self-operated-with-the-labour-counted).

Record those as decisions rather than as things that survived a review. A cost with a stated
justification is defensible next time; one that merely was not cut this round will be cut next
round, and the reliability or performance it was buying goes with it.

This is the mechanism that protects the other pillars from cost pressure. [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) and [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/)
produce the targets; this best practice is where those targets convert into a defence of specific
spending.

**On STACKIT.** No platform feature applies. What helps is that the justification can be attached
to a project rather than to a resource, since the project is the unit the <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost
Dashboard</LinkChip> reports on. A
project whose stated purpose includes an availability target is easier to defend than a line item.

**Tradeoffs.** **Cost Optimization**, against itself, deliberately. This best practice exists to
stop the pillar from optimizing itself into a workload that fails its other requirements.

**Verify.** In your last cost review, was any cost increased or explicitly protected? If every
conclusion was a reduction, what was the review optimizing for?

---

## COST 7.4 Reconcile with the reliability and performance decisions rather than around them

**Risk if not established:** High

The most damaging cost reductions are the ones made without the other pillars in the room. Backup
retention shortened, a standby downsized, a third replica dropped, a log retention halved. Each is
individually defensible and none is announced as a reduction in reliability.

The result is a workload whose actual recovery capability has drifted away from its stated targets
without any decision having been made. [`REL 1.4`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-14-have-each-target-agreed-and-recorded-by-someone-accountable-for-the-outcome) and the
[Cost Optimization tradeoffs](/architecture/pillars/cost-optimization/tradeoffs/) both describe this, and this is where it is
prevented.

The mechanism is that a change breaching a stated target becomes a decision somebody has to make
explicitly. That only works if the targets exist, which is the dependency this question has on
[`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) and [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/).

Do the reviews together rather than sequentially. The cost review, the sizing review under
[`PERF 9.1`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/#perf-91-re-examine-sizing-against-measured-demand-on-a-cadence) and the reliability targets look at the same components and reach conclusions that need
reconciling. Reconciling them in one conversation is cheaper than discovering the conflict after
one of them has acted.

**On STACKIT.** Mark the costs that exist because of a compliance obligation, since those look
like technical choices from a spreadsheet and are not. [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) retention, [`SOV 4`](/architecture/pillars/sovereignty/sov-04-key-ownership/) key management
and [`SOV 5`](/architecture/pillars/sovereignty/sov-05-operator-access/) confidential computing all carry cost that is not discretionary, and the [Sovereignty
tradeoffs](/architecture/pillars/sovereignty/tradeoffs/) explain why they are cut by people who do not know
what they are cutting.

**Tradeoffs.** **Operational Excellence.** A joint review is harder to schedule and involves more
people. The alternative is a sequence of locally rational decisions that are collectively wrong.

**Verify.** For the last three cost reductions you made, which stated target did each one affect?
Who agreed that the reduced level was acceptable?

---

## Related

- [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume) Flow ranking, which this question reuses
- [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) Attribution, without which the flow view cannot be assembled
- [`COST 5`](/architecture/pillars/cost-optimization/cost-05-environments/) and [`COST 6`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/), where the poorest value-to-cost ratios usually sit
- [`REL 1.4`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-14-have-each-target-agreed-and-recorded-by-someone-accountable-for-the-outcome) and [`PERF 1.3`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/#perf-13-have-it-agreed-by-someone-who-represents-the-users), whose agreed targets make protection possible
- <LinkChip href="/architecture/pillars/cost-optimization/tradeoffs/">Cost Optimization tradeoffs</LinkChip>
