Zum Inhalt springen
Beta

COST 1. How do you estimate what a design will cost before you build it?

Zuletzt aktualisiert am

The decisions that determine what a workload costs are made early: the topology, the data model, whether state is replicated synchronously, whether anything can scale to zero, and whether you operate a component or consume it. By the time there is a bill, those are settled and the remaining levers are the small ones.

That is why cost optimization done as a later project disappoints. It finds the oversized instances and the forgotten volumes, which is worth having and rarely transformative, while the structural cost sits behind a rewrite nobody will fund.

  • COST 1.1 Price the design while it is still cheap to change
  • COST 1.2 Include operational labour, not only resource consumption
  • COST 1.3 State what drives the number, so later divergence is diagnosable
  • COST 1.4 Model per environment rather than once for the workload

COST 1.1 Price the design while it is still cheap to change

Section titled “COST 1.1 Price the design while it is still cheap to change”

Risk if not established: Medium

Treat a cost estimate as part of a design review, alongside the availability target from REL 1 and the performance target from PERF 1. All three are requirements, and a design that satisfies two of them is not finished.

Precision is not the point. An estimate within a factor of two is enough to reveal that one component dominates, that a topology costs four times an alternative, or that the design is an order of magnitude away from what the business expects. Those are the findings that change a design, and none of them need a precise number.

Price the alternatives rather than the chosen design alone. The useful output is a comparison: this topology against that one, managed against self-operated, synchronous replication against asynchronous. A single number tells you what it costs and not whether it is reasonable.

Include the things that are easy to forget because they are not compute: storage that accumulates, traffic between zones, backups and their retention, telemetry volume, and the environments beyond production.

On STACKIT. Pricing dimensions differ per service and are the input to any estimate: machine type variants for Compute Engine, flavors and performance classes for managed databases, service plans elsewhere. Which dimension dominates is service-specific and is worth checking rather than assuming compute dominates, because for data-heavy workloads it frequently does not.

The Cost Dashboard is the other half of the loop. It supplies what the design actually cost once it exists, which is what COST 9 compares against the estimate.

Tradeoffs. Operational Excellence. Costs design time, and the estimate will be wrong. It is still the only way to discover that a design is unaffordable while changing it is cheap.

Verify. For your most recent significant design, what was it estimated to cost before it was built? How close was that to reality?


COST 1.2 Include operational labour, not only resource consumption

Section titled “COST 1.2 Include operational labour, not only resource consumption”

Risk if not established: Medium

A model counting only what the platform bills systematically favours self-operated components, because their largest cost is people and people are not on the invoice.

Count the labour: building it, patching it under SEC 8.2, monitoring it, carrying the pager for it under OPS 1.1, and the expertise that has to exist somewhere in the team. Those are continuing costs and they scale with the number of components rather than with their size.

This is the comparison that makes a managed database look expensive next to a virtual machine and frequently cheaper once somebody counts the engineer-hours. It is also the comparison most cost models omit, which is why the conclusion so often runs the wrong way.

Include the opportunity cost where it is large. Engineering time spent operating a component is time not spent on the product, and for a small team that is the binding constraint rather than the budget.

On STACKIT. The choice appears at every layer: a database on Compute Engine against a managed one, a self-run CI system against STACKIT Pipelines , your own backup automation against Server Backup Management .

In each case the managed option costs more per unit and removes work. Whether that is a good trade depends on how much the work costs you, which is a number your organization has and the platform does not.

Tradeoffs. Sovereignty & Compliance. A managed service means the provider operates it, which is the shared responsibility split SOV 8 asks you to map rather than a problem, and it is worth being explicit about in a model that is otherwise purely financial.

Verify. For a component you operate yourself, how many engineer-hours per month does it consume? Was that figure in the comparison when the decision was made?


COST 1.3 State what drives the number, so later divergence is diagnosable

Section titled “COST 1.3 State what drives the number, so later divergence is diagnosable”

Risk if not established: Medium

An estimate that is a single figure can only be right or wrong. An estimate that states its drivers can be diagnosed when reality diverges, which is what COST 9 needs to be useful.

Write down the assumptions the number rests on: expected request volume, data growth per month, average object size, retention periods, the number of environments, and the peak-to-average ratio. Each is a quantity that will turn out differently, and knowing which one moved is the difference between “costs are higher than expected” and “data grew three times faster than modelled”.

Identify which driver dominates. Most workloads have one or two costs that account for most of the bill, and the sensitivity of the model to those is what matters. A ten percent error on the dominant driver outweighs a factor-of-two error on everything else.

Model the growth rather than the starting point. A workload that is affordable today and whose cost scales linearly with a data set growing monthly has a date attached to it, and that date is worth knowing before it arrives.

On STACKIT. The drivers are service-specific and stated in each service’s pricing dimensions. What is worth checking early is which dimension your workload actually loads: a database sized for capacity but limited by I/O is paying on one axis and constrained on another, which PERF 3.3 also addresses.

Tradeoffs. Little. Writing down assumptions costs minutes and is the difference between a model that can be corrected and one that can only be replaced.

Verify. For your cost model, what are the three largest drivers and what value was assumed for each? Which of those has moved most since?


COST 1.4 Model per environment rather than once for the workload

Section titled “COST 1.4 Model per environment rather than once for the workload”

Risk if not established: Medium

A model that prices production and assumes the rest is a rounding error is usually wrong, because non-production environments frequently cost a substantial share of the total and nobody has ever looked.

Price each environment for what it is actually for, which is COST 5. A test environment mirroring production topology costs close to production, and if that is the design it belongs in the model rather than arriving as a surprise.

Count them all, including the ones that are not on the diagram: the demo environment, the one for a migration that finished, the personal sandboxes. SEC 1.4 finds the same list for a different reason.

Model their lifetime as well as their size. An environment that exists for two weeks per quarter costs a fraction of one that runs continuously, and whether it can be created on demand is a design decision that OPS 3 makes possible.

On STACKIT. Environment separation is expressed as separate projects, which is the same structure SEC 2.2 and COST 2 need. That makes per-environment cost visible in the Cost Dashboard without extra work, since it breaks costs down per project.

That only holds if the project structure separates environments. Where several environments share a project, their costs are combined and no later analysis can separate them, which is one of several reasons the hierarchy decision in COST 2.1 is worth making deliberately.

Tradeoffs. Operational Excellence. More projects to manage, which is the same cost segmentation carries in SEC 2.2.

Verify. How many environments does your workload have, and what does each cost per month? Which of those figures did you have to estimate rather than look up?


  • COST 2 Attribution, which makes the actual figures available per owner
  • COST 9 Review cadence, which compares reality against this model
  • COST 5 Environments, priced here and sized there
  • REL 1 and PERF 1, the other two requirements a design has to satisfy
  • OPS 1.4 Funding operational work, which this model should make visible