---
id: COST01
pillar: cost-optimization
title: COST 1. How do you estimate what a design will cost before you build it?
description: By the time the first invoice arrives, most of the cost has already been decided. A design nobody priced carries a requirement that nobody ever examined.
status: draft
services: []
sidebar:
  order: 10
  label: Cost model
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/cost-01-cost-model/"
source_file: "docs/architecture/pillars/cost-optimization/cost-01-cost-model.mdx"
---

The decisions that determine what a workload costs are made early: the topology, the data model,
whether state is replicated synchronously, whether anything can scale to zero, and whether you
operate a component or consume it. By the time there is a bill, those are settled and the
remaining levers are the small ones.

That is why cost optimization done as a later project disappoints. It finds the oversized
instances and the forgotten volumes, which is worth having and rarely transformative, while the
structural cost sits behind a rewrite nobody will fund.

## Best practices

- [`COST 1.1`](/architecture/pillars/cost-optimization/cost-01-cost-model/#cost-11-price-the-design-while-it-is-still-cheap-to-change) Price the design while it is still cheap to change
- [`COST 1.2`](/architecture/pillars/cost-optimization/cost-01-cost-model/#cost-12-include-operational-labour-not-only-resource-consumption) Include operational labour, not only resource consumption
- [`COST 1.3`](/architecture/pillars/cost-optimization/cost-01-cost-model/#cost-13-state-what-drives-the-number-so-later-divergence-is-diagnosable) State what drives the number, so later divergence is diagnosable
- [`COST 1.4`](/architecture/pillars/cost-optimization/cost-01-cost-model/#cost-14-model-per-environment-rather-than-once-for-the-workload) Model per environment rather than once for the workload

---

## COST 1.1 Price the design while it is still cheap to change

**Risk if not established:** Medium

Treat a cost estimate as part of a design review, alongside the availability target from [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/)
and the performance target from [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/). All three are requirements, and a design that satisfies
two of them is not finished.

Precision is not the point. An estimate within a factor of two is enough to reveal that one
component dominates, that a topology costs four times an alternative, or that the design is an
order of magnitude away from what the business expects. Those are the findings that change a
design, and none of them need a precise number.

Price the alternatives rather than the chosen design alone. The useful output is a comparison:
this topology against that one, managed against self-operated, synchronous replication against
asynchronous. A single number tells you what it costs and not whether it is reasonable.

Include the things that are easy to forget because they are not compute: storage that accumulates,
traffic between zones, backups and their retention, telemetry volume, and the environments beyond
production.

**On STACKIT.** Pricing dimensions differ per service and are the input to any estimate: machine
type variants for Compute Engine, flavors and performance classes for managed databases, service
plans elsewhere. Which dimension dominates is service-specific and is worth checking rather than
assuming compute dominates, because for data-heavy workloads it frequently does not.

The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost Dashboard</LinkChip> is the
other half of the loop. It supplies what the design actually cost once it exists, which is what
[`COST 9`](/architecture/pillars/cost-optimization/cost-09-review-cadence/) compares against the estimate.

**Tradeoffs.** **Operational Excellence.** Costs design time, and the estimate will be wrong. It is
still the only way to discover that a design is unaffordable while changing it is cheap.

**Verify.** For your most recent significant design, what was it estimated to cost before it was
built? How close was that to reality?

---

## COST 1.2 Include operational labour, not only resource consumption

**Risk if not established:** Medium

A model counting only what the platform bills systematically favours self-operated components,
because their largest cost is people and people are not on the invoice.

Count the labour: building it, patching it under [`SEC 8.2`](/architecture/pillars/security/sec-08-hardening-and-patching/#sec-82-define-a-patch-cadence-with-a-maximum-exposure-window-per-severity), monitoring it, carrying the pager for
it under [`OPS 1.1`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-11-give-the-team-that-builds-a-system-a-genuine-stake-in-running-it), and the expertise that has to exist somewhere in the team. Those are
continuing costs and they scale with the number of components rather than with their size.

This is the comparison that makes a managed database look expensive next to a virtual machine and
frequently cheaper once somebody counts the engineer-hours. It is also the comparison most cost
models omit, which is why the conclusion so often runs the wrong way.

Include the opportunity cost where it is large. Engineering time spent operating a component is
time not spent on the product, and for a small team that is the binding constraint rather than the
budget.

**On STACKIT.** The choice appears at every layer: a database on Compute Engine against a managed
one, a self-run CI system against <LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/basics/stackit-pipelines/">STACKIT
Pipelines</LinkChip>,
your own backup automation against <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/">Server Backup
Management</LinkChip>.

In each case the managed option costs more per unit and removes work. Whether that is a good trade
depends on how much the work costs you, which is a number your organization has and the platform
does not.

**Tradeoffs.** **Sovereignty & Compliance.** A managed service means the provider operates it,
which is the shared responsibility split [`SOV 8`](/architecture/pillars/sovereignty/sov-08-compliance-mapping/) asks you to map rather than a problem, and it is
worth being explicit about in a model that is otherwise purely financial.

**Verify.** For a component you operate yourself, how many engineer-hours per month does it
consume? Was that figure in the comparison when the decision was made?

---

## COST 1.3 State what drives the number, so later divergence is diagnosable

**Risk if not established:** Medium

An estimate that is a single figure can only be right or wrong. An estimate that states its
drivers can be diagnosed when reality diverges, which is what [`COST 9`](/architecture/pillars/cost-optimization/cost-09-review-cadence/) needs to be useful.

Write down the assumptions the number rests on: expected request volume, data growth per month,
average object size, retention periods, the number of environments, and the peak-to-average ratio.
Each is a quantity that will turn out differently, and knowing which one moved is the difference
between "costs are higher than expected" and "data grew three times faster than modelled".

Identify which driver dominates. Most workloads have one or two costs that account for most of the
bill, and the sensitivity of the model to those is what matters. A ten percent error on the
dominant driver outweighs a factor-of-two error on everything else.

Model the growth rather than the starting point. A workload that is affordable today and whose
cost scales linearly with a data set growing monthly has a date attached to it, and that date is
worth knowing before it arrives.

**On STACKIT.** The drivers are service-specific and stated in each service's pricing dimensions.
What is worth checking early is which dimension your workload actually loads: a database sized for
capacity but limited by I/O is paying on one axis and constrained on another, which [`PERF 3.3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-33-match-the-resource-shape-to-the-workload-shape)
also addresses.

**Tradeoffs.** Little. Writing down assumptions costs minutes and is the difference between a
model that can be corrected and one that can only be replaced.

**Verify.** For your cost model, what are the three largest drivers and what value was assumed for
each? Which of those has moved most since?

---

## COST 1.4 Model per environment rather than once for the workload

**Risk if not established:** Medium

A model that prices production and assumes the rest is a rounding error is usually wrong, because
non-production environments frequently cost a substantial share of the total and nobody has ever
looked.

Price each environment for what it is actually for, which is [`COST 5`](/architecture/pillars/cost-optimization/cost-05-environments/). A test environment
mirroring production topology costs close to production, and if that is the design it belongs in
the model rather than arriving as a surprise.

Count them all, including the ones that are not on the diagram: the demo environment, the one for
a migration that finished, the personal sandboxes. [`SEC 1.4`](/architecture/pillars/security/sec-01-security-baseline/#sec-14-cover-the-whole-estate-including-the-parts-nobody-claims) finds the same list for a different
reason.

Model their lifetime as well as their size. An environment that exists for two weeks per quarter
costs a fraction of one that runs continuously, and whether it can be created on demand is a
design decision that [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) makes possible.

**On STACKIT.** Environment separation is expressed as separate projects, which is the same
structure [`SEC 2.2`](/architecture/pillars/security/sec-02-segmentation/#sec-22-structure-the-resource-hierarchy-so-it-can-carry-the-boundaries-you-need) and [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) need. That makes per-environment cost visible in the <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost
Dashboard</LinkChip> without extra
work, since it breaks costs down per project.

That only holds if the project structure separates environments. Where several environments share
a project, their costs are combined and no later analysis can separate them, which is one of
several reasons the hierarchy decision in [`COST 2.1`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-21-make-the-resource-hierarchy-reflect-who-pays) is worth making deliberately.

**Tradeoffs.** **Operational Excellence.** More projects to manage, which is the same cost
segmentation carries in [`SEC 2.2`](/architecture/pillars/security/sec-02-segmentation/#sec-22-structure-the-resource-hierarchy-so-it-can-carry-the-boundaries-you-need).

**Verify.** How many environments does your workload have, and what does each cost per month?
Which of those figures did you have to estimate rather than look up?

---

## Related

- [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) Attribution, which makes the actual figures available per owner
- [`COST 9`](/architecture/pillars/cost-optimization/cost-09-review-cadence/) Review cadence, which compares reality against this model
- [`COST 5`](/architecture/pillars/cost-optimization/cost-05-environments/) Environments, priced here and sized there
- [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) and [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/), the other two requirements a design has to satisfy
- [`OPS 1.4`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-14-fund-operational-work-explicitly-rather-than-expecting-it-to-fit-in-the-gaps) Funding operational work, which this model should make visible
