Cost Optimization
Last updated on
Is the money buying business value?
Cost optimization is not cost reduction. A workload can be made arbitrarily cheap by making it useless, and the cheapest architecture is almost never the right one. The question this pillar asks is narrower and harder: for every unit of spend, is it buying something the business wants?
That reframing matters because the two goals lead to different work. Cost reduction is a project, usually launched under pressure, that finds the obvious waste and then stops. Cost optimization is a continuous practice that keeps spend aligned with value as both drift, and they always drift, because requirements change, traffic patterns change, and yesterday’s correct instance size is today’s waste without anyone having done anything wrong.
What this pillar covers
Section titled “What this pillar covers”- Building a cost model before building the workload
- Attribution: knowing which team, product, or flow caused a cost
- Right-sizing as a continuous activity rather than an event
- Choosing services and tiers that match the actual requirement
- Environments: non-production is where the quiet waste accumulates
- Data lifecycle costs, which grow whether you attend to them or not
- Directing spend by business value rather than by resource
- Visibility, anomaly detection, and a review cadence
What it does not cover
Section titled “What it does not cover”Efficiency in the sense of doing more work per unit of resource belongs to Performance Efficiency, the two pillars agree far more often than they conflict, and a performance optimization is usually also a cost optimization. Reducing resource consumption for environmental reasons belongs to Sustainability, which mostly agrees with this pillar and diverges in specific places worth knowing about.
The cost of not being reliable, secure, or compliant is real and is not modelled here; it appears in the tradeoffs of the pillars that address those risks.
The central idea
Section titled “The central idea”Cost you cannot attribute is cost you cannot manage. This is the question everything else depends on, and the one most often deferred because it requires organizational work rather than technical work.
An invoice that says the platform cost a certain amount last month is not actionable. An invoice that says which product, which environment, and which team caused each part of it turns a finance conversation into an engineering one, and puts the information in front of the people who can act on it, which is the only place it produces change.
The second idea: the cheapest resource is the one you do not run. Optimization instinctively reaches for cheaper tiers and better discounts, which are real but bounded. The larger wins are structurally different, deleting environments nobody uses, shutting down capacity outside working hours, removing a caching layer that a query fix made unnecessary. Elimination beats discounting, and it is where the first pass should look.
Where to start
Section titled “Where to start”Questions
Section titled “Questions”Nine questions. Numbers follow the order the decisions are usually made in and do not indicate priority.
COST 1: How do you estimate what a design will cost before you build it?
Section titled “COST 1: How do you estimate what a design will cost before you build it?”Estimate what the design will cost, per environment, before it is built, and state what drives the number, so that later divergence is diagnosable rather than merely surprising. A design nobody priced has an unexamined requirement.
COST 2: How do you attribute cost to the team or product that causes it?
Section titled “COST 2: How do you attribute cost to the team or product that causes it?”Structure the resource hierarchy and labelling so that every cost traces to a product, an environment, and a team. Attribution is a prerequisite for everything else in this pillar and is expensive to retrofit.
COST 3: How do you keep provisioned capacity matched to measured demand?
Section titled “COST 3: How do you keep provisioned capacity matched to measured demand?”Compare provisioned capacity against measured usage on a cadence, and act on the gap. Sizing decisions decay; a system nobody has re-examined is a system that is now approximately the wrong size.
COST 4: How do you choose the service and tier that matches the requirement?
Section titled “COST 4: How do you choose the service and tier that matches the requirement?”Select the cheapest option that meets the stated requirement rather than the one that meets the requirement you might have later. Re-examine when the requirement changes, tier decisions are rarely revisited even when their justification has disappeared.
COST 5: How do you size non-production environments for their real purpose?
Section titled “COST 5: How do you size non-production environments for their real purpose?”A test environment does not need production topology, production redundancy, or production uptime. Size each environment for what it is actually used for, and shut it down when it is not being used.
COST 6: How do you manage the cost of data across its whole lifecycle?
Section titled “COST 6: How do you manage the cost of data across its whole lifecycle?”Data accumulates by default and is almost never deleted by default. Define retention per data set, tier or archive data as it cools, and delete what has no further use, including old backups, snapshots, and logs.
COST 7: How do you direct cost scrutiny by business value?
Section titled “COST 7: How do you direct cost scrutiny by business value?”Direct cost scrutiny by business value, using the flow ranking from REL 2. Optimizing a resource
list top-down by price finds the largest line items, which is not the same as finding the spend
that buys the least.
COST 8: How do you make cost visible to the people who cause it?
Section titled “COST 8: How do you make cost visible to the people who cause it?”Put current and trending cost in front of the engineers whose decisions cause it, and alert on unexpected increases quickly enough to act. A monthly invoice is a report; a same-week anomaly alert is a control.
COST 9: How do you review actual spend against the model on a fixed cadence?
Section titled “COST 9: How do you review actual spend against the model on a fixed cadence?”Compare reality against COST 1’s model on a defined schedule, with a named owner, and act on the
delta, either by fixing the workload or by correcting the model. Without a cadence, optimization
happens under budget pressure, which is the most expensive time to do it.
Related
Section titled “Related”- Design principles
- Tradeoffs
- Sustainability: overlapping levers, different objective