COST 1. How do you estimate what a design will cost before you build it?
Last updated on
The decisions that determine what a workload costs are made early: the topology, the data model, whether state is replicated synchronously, whether anything can scale to zero, and whether you operate a component or consume it. By the time there is a bill, those are settled and the remaining levers are the small ones.
That is why cost optimization done as a later project disappoints. It finds the oversized instances and the forgotten volumes, which is worth having and rarely transformative, while the structural cost sits behind a rewrite nobody will fund.
Best practices
Section titled “Best practices”COST 1.1Price the design while it is still cheap to changeCOST 1.2Include operational labour, not only resource consumptionCOST 1.3State what drives the number, so later divergence is diagnosableCOST 1.4Model per environment rather than once for the workload
COST 1.1 Price the design while it is still cheap to change
Section titled “COST 1.1 Price the design while it is still cheap to change”Risk if not established: Medium
Treat a cost estimate as part of a design review, alongside the availability target from REL 1
and the performance target from PERF 1. All three are requirements, and a design that satisfies
two of them is not finished.
Precision is not the point. An estimate within a factor of two is enough to reveal that one component dominates, that a topology costs four times an alternative, or that the design is an order of magnitude away from what the business expects. Those are the findings that change a design, and none of them need a precise number.
Price the alternatives rather than the chosen design alone. The useful output is a comparison: this topology against that one, managed against self-operated, synchronous replication against asynchronous. A single number tells you what it costs and not whether it is reasonable.
Include the things that are easy to forget because they are not compute: storage that accumulates, traffic between zones, backups and their retention, telemetry volume, and the environments beyond production.
On STACKIT. Pricing dimensions differ per service and are the input to any estimate: machine type variants for Compute Engine, flavors and performance classes for managed databases, service plans elsewhere. Which dimension dominates is service-specific and is worth checking rather than assuming compute dominates, because for data-heavy workloads it frequently does not.
The Cost Dashboard is the
other half of the loop. It supplies what the design actually cost once it exists, which is what
COST 9 compares against the estimate.
Tradeoffs. Operational Excellence. Costs design time, and the estimate will be wrong. It is still the only way to discover that a design is unaffordable while changing it is cheap.
Verify. For your most recent significant design, what was it estimated to cost before it was built? How close was that to reality?
COST 1.2 Include operational labour, not only resource consumption
Section titled “COST 1.2 Include operational labour, not only resource consumption”Risk if not established: Medium
A model counting only what the platform bills systematically favours self-operated components, because their largest cost is people and people are not on the invoice.
Count the labour: building it, patching it under SEC 8.2, monitoring it, carrying the pager for
it under OPS 1.1, and the expertise that has to exist somewhere in the team. Those are
continuing costs and they scale with the number of components rather than with their size.
This is the comparison that makes a managed database look expensive next to a virtual machine and frequently cheaper once somebody counts the engineer-hours. It is also the comparison most cost models omit, which is why the conclusion so often runs the wrong way.
Include the opportunity cost where it is large. Engineering time spent operating a component is time not spent on the product, and for a small team that is the binding constraint rather than the budget.
On STACKIT. The choice appears at every layer: a database on Compute Engine against a managed one, a self-run CI system against STACKIT Pipelines , your own backup automation against Server Backup Management .
In each case the managed option costs more per unit and removes work. Whether that is a good trade depends on how much the work costs you, which is a number your organization has and the platform does not.
Tradeoffs. Sovereignty & Compliance. A managed service means the provider operates it,
which is the shared responsibility split SOV 8 asks you to map rather than a problem, and it is
worth being explicit about in a model that is otherwise purely financial.
Verify. For a component you operate yourself, how many engineer-hours per month does it consume? Was that figure in the comparison when the decision was made?
COST 1.3 State what drives the number, so later divergence is diagnosable
Section titled “COST 1.3 State what drives the number, so later divergence is diagnosable”Risk if not established: Medium
An estimate that is a single figure can only be right or wrong. An estimate that states its
drivers can be diagnosed when reality diverges, which is what COST 9 needs to be useful.
Write down the assumptions the number rests on: expected request volume, data growth per month, average object size, retention periods, the number of environments, and the peak-to-average ratio. Each is a quantity that will turn out differently, and knowing which one moved is the difference between “costs are higher than expected” and “data grew three times faster than modelled”.
Identify which driver dominates. Most workloads have one or two costs that account for most of the bill, and the sensitivity of the model to those is what matters. A ten percent error on the dominant driver outweighs a factor-of-two error on everything else.
Model the growth rather than the starting point. A workload that is affordable today and whose cost scales linearly with a data set growing monthly has a date attached to it, and that date is worth knowing before it arrives.
On STACKIT. The drivers are service-specific and stated in each service’s pricing dimensions.
What is worth checking early is which dimension your workload actually loads: a database sized for
capacity but limited by I/O is paying on one axis and constrained on another, which PERF 3.3
also addresses.
Tradeoffs. Little. Writing down assumptions costs minutes and is the difference between a model that can be corrected and one that can only be replaced.
Verify. For your cost model, what are the three largest drivers and what value was assumed for each? Which of those has moved most since?
COST 1.4 Model per environment rather than once for the workload
Section titled “COST 1.4 Model per environment rather than once for the workload”Risk if not established: Medium
A model that prices production and assumes the rest is a rounding error is usually wrong, because non-production environments frequently cost a substantial share of the total and nobody has ever looked.
Price each environment for what it is actually for, which is COST 5. A test environment
mirroring production topology costs close to production, and if that is the design it belongs in
the model rather than arriving as a surprise.
Count them all, including the ones that are not on the diagram: the demo environment, the one for
a migration that finished, the personal sandboxes. SEC 1.4 finds the same list for a different
reason.
Model their lifetime as well as their size. An environment that exists for two weeks per quarter
costs a fraction of one that runs continuously, and whether it can be created on demand is a
design decision that OPS 3 makes possible.
On STACKIT. Environment separation is expressed as separate projects, which is the same
structure SEC 2.2 and COST 2 need. That makes per-environment cost visible in the Cost
Dashboard without extra
work, since it breaks costs down per project.
That only holds if the project structure separates environments. Where several environments share
a project, their costs are combined and no later analysis can separate them, which is one of
several reasons the hierarchy decision in COST 2.1 is worth making deliberately.
Tradeoffs. Operational Excellence. More projects to manage, which is the same cost
segmentation carries in SEC 2.2.
Verify. How many environments does your workload have, and what does each cost per month? Which of those figures did you have to estimate rather than look up?
Related
Section titled “Related”COST 2Attribution, which makes the actual figures available per ownerCOST 9Review cadence, which compares reality against this modelCOST 5Environments, priced here and sized thereREL 1andPERF 1, the other two requirements a design has to satisfyOPS 1.4Funding operational work, which this model should make visible