COST 4. How do you choose the service and tier that matches the requirement?
Zuletzt aktualisiert am
Two mistakes account for most of the cost in this question, and they run in opposite directions. Choosing a service because it is familiar rather than because it fits, and choosing a tier that meets a requirement nobody has stated.
Both are made once, early, and inherited. A tier chosen for a peak that never arrived costs the difference every month for years, and nothing draws attention to it.
Best practices
Section titled “Best practices”COST 4.1Choose the cheapest option that meets the stated requirementCOST 4.2Compare managed against self-operated with the labour countedCOST 4.3Know what actually drives the bill for each optionCOST 4.4Re-examine when the requirement changes
COST 4.1 Choose the cheapest option that meets the stated requirement
Section titled “COST 4.1 Choose the cheapest option that meets the stated requirement”Risk if not established: Medium
The requirement has to exist first, which is why this question depends on PERF 1 and REL 1
having produced numbers. Without them, tier selection is a judgement about how much reliability
and speed feels right, and that judgement is consistently generous.
Choose against the stated requirement rather than against a possible future one. Provisioning for
a requirement you might have later means paying for it from today, and the alternative is usually
available: choose for now and re-examine under COST 4.4.
Where the difference between tiers is small, take the higher one and stop thinking about it. Where it is a factor of several, it warrants the analysis. Spending an afternoon to save a few euros a month is its own kind of waste.
Watch for tiers chosen for a property that is not actually used. A high-availability configuration for a workload whose stated RTO is a day, or a performance class chosen for headroom on an axis the workload never loads.
On STACKIT. Options are structured differently per service and the dimensions are worth reading before choosing: machine type variants for Compute Engine, flavors and performance classes for managed databases, and service plans elsewhere.
The single-instance against replica-set choice for a managed database is the one with the largest
cost difference and the clearest requirement behind it. REL 1 produces the target that decides
it, and the planning
guidance
names the three-replica type for production use, which is a recommendation rather than a
requirement your particular workload necessarily has.
Tradeoffs. Reliability. The cheapest option that meets the requirement has no margin for
the requirement being wrong, which is why REL 1.1 insists the requirement come from the business
rather than from a preference.
Verify. For each service you use, which tier is selected and which stated requirement justifies it? How many were chosen without a stated requirement?
COST 4.2 Compare managed against self-operated with the labour counted
Section titled “COST 4.2 Compare managed against self-operated with the labour counted”Risk if not established: Medium
A comparison of the invoice line alone reliably concludes that self-operated is cheaper, because its dominant cost is people and people are not on the invoice.
Count what COST 1.2 describes: building it, patching it, monitoring it, backing it up, carrying
the pager, and the expertise that has to exist. Then add the failure modes you inherit, since a
self-operated database has an availability that depends on your operations rather than on a
published commitment.
The comparison usually reverses once labour is counted, and not always. Self-operating makes sense where the managed option does not fit the requirement, where you need a version or configuration the service does not offer, or where the scale is large enough that the per-unit premium exceeds the labour.
Include the reversibility. Moving from managed to self-operated later is possible; the reverse usually is too. What is harder to reverse is the expertise decision, because a team that has never operated the component cannot start quickly.
On STACKIT. The choice exists at most layers. A database on Compute Engine against PostgreSQL Flex , your own CI against STACKIT Pipelines , your own backup automation against Server Backup Management , your own scheduler against Automation Service .
One asymmetry is worth stating: a managed service comes with a published availability commitment
in its service certificate , which is an input
to the composition in REL 1.3. A self-operated equivalent has whatever availability your
operations achieve, and that number is not published anywhere because nobody has measured it.
Tradeoffs. Sovereignty & Compliance. Managed means the provider operates it, which changes
the shared responsibility split that SOV 8 maps. That is a factor rather than an objection and
it belongs in the comparison.
Verify. For each self-operated component, what would the managed equivalent cost, and how many engineer-hours per month does the current arrangement consume?
COST 4.3 Know what actually drives the bill for each option
Section titled “COST 4.3 Know what actually drives the bill for each option”Risk if not established: Medium
Cost intuition transfers badly between services. The dimension that dominates one is negligible in another, and a design optimized for the wrong dimension saves nothing.
For each service on a ranked flow, establish which dimension carries the cost. Sometimes it is capacity, sometimes throughput, sometimes the number of operations, sometimes data transfer, sometimes the number of instances regardless of their size.
The ones that surprise people are the operation-count and transfer dimensions, because they scale with behaviour rather than with size. A storage bucket holding very little data and receiving millions of small requests can cost more than one holding a great deal and receiving few.
Knowing the dominant dimension is what makes an optimization worth doing. Compressing objects
reduces a capacity bill and does nothing for a request-count bill; batching does the reverse.
PERF 6.3 describes the same techniques from the performance side and they pay differently here.
On STACKIT. Object Storage is the clearest case where behaviour rather than volume drives both cost and performance. The performance guidance covers object size, request rate and parallelism, and those are the same dimensions to think about when the bill is the concern.
Inter-zone traffic is the dimension most often forgotten in a multi-zone design, and it is a
direct consequence of the redundancy chosen under REL 4.1. A chatty component distributed across
zones pays for every call, which is a cost that did not exist before the topology changed.
Tradeoffs. Little. This is an analysis that changes which optimizations are worth attempting, and getting it wrong means effort spent on the dimension that was not the bill.
Verify. For your three largest cost items, which dimension drives each? Which of those did you have to look up rather than know?
COST 4.4 Re-examine when the requirement changes
Section titled “COST 4.4 Re-examine when the requirement changes”Risk if not established: Medium
Tier decisions are among the stickiest in an estate. They are made at design time, they work, and nothing about a working configuration draws attention to itself.
Requirements change underneath them. A workload that was customer-facing becomes internal. A reporting database that needed to be fast now runs overnight. A service that carried regulated data no longer does. Each of those changes the tier that is justified, and none of them triggers a review.
Tie the re-examination to the change rather than to a calendar. A change in classification under
SEC 3, a change in the flow ranking under REL 2.2, or a change in the stated target under
REL 1 and PERF 1 should each prompt the question of whether the tier still fits.
Look in both directions. The reflex is to check whether a tier is now too small; the saving is usually in the ones that are now too large.
On STACKIT. Whether the re-examination can act on its conclusion depends on the reversibility
sort from COST 3.3. A service plan that can be changed is worth reviewing frequently; one that
requires a migration is worth reviewing when something else already justifies the disruption, such
as a version upgrade under OPS 6.3.
Tradeoffs. Operational Excellence. Acting on a re-examination means a change with its own risk, which is why the reversibility sort decides whether it is worth acting on rather than recording.
Verify. Which of your tier decisions were made more than a year ago? For each, has the requirement that justified it changed?
Related
Section titled “Related”COST 1.2Cost model, where the labour comparison belongsCOST 3.3Reversibility, which decides whether a re-examination can actPERF 3Selection and sizing, the same decisions serving the performance targetREL 1.3Availability composition, which service certificates feedSUS 2Right-sizing, which reaches similar conclusions from consumption