Zum Inhalt springen
Beta

Cost Optimization: design principles

Zuletzt aktualisiert am

1. Cost is a design constraint, not a cleanup task

Section titled “1. Cost is a design constraint, not a cleanup task”

The decisions that determine what a workload costs are made early: the data model, the topology, the choice between managed and self-operated, whether state is replicated synchronously, whether the system can scale to zero. By the time the first invoice arrives, most of the cost has already been decided and the remaining levers are the small ones.

Which is why cost optimization done as a retroactive project produces disappointing results. It finds the oversized instances and the forgotten volumes, worth having, rarely transformative, while the structural costs are untouchable without a rewrite nobody will fund.

Treat a cost estimate as part of a design review, alongside the availability target and the threat model. A design whose cost nobody estimated is a design with an unexamined requirement.

Not all parts of a workload deserve equal investment. The flow that processes revenue and the internal admin page that three people use once a month have different claims on the budget, and an architecture that treats them identically is wrong in one direction for one of them.

This is the same ranking that [REL 2](/architecture/pillars/reliability/rel-02-critical-flows/) produces, used for a different purpose, which is why doing it once serves two pillars. Once flows are ranked by business impact, both reliability spend and cost scrutiny can follow the ranking instead of being applied uniformly.

The uncomfortable half: this principle also justifies spending more in places. Cost optimization that only ever reduces is not following value. It is following a target.

3. Cost you cannot attribute is cost you cannot manage

Section titled “3. Cost you cannot attribute is cost you cannot manage”

An aggregate number changes nobody’s behaviour. The engineer who could turn off the idle cluster does not see the invoice, and the person who sees the invoice does not know which cluster is idle.

Attribution closes that gap, and it is structural rather than analytical: it depends on the resource hierarchy and labelling reflecting who actually owns what. Retrofitting attribution onto an environment that grew without it is genuinely hard, which is why COST 2 sits near the front.

The test is whether you can answer, without a project, what a given product cost last month across all its environments. If that takes a week of spreadsheet work, cost is not being managed; it is being reported.

4. The cheapest resource is the one you do not run

Section titled “4. The cheapest resource is the one you do not run”

Optimization reflexively looks for cheaper, smaller instances, lower tiers, better commitments. Those are real and bounded. The unbounded wins come from elimination.

The development environment nobody has opened in eight months. The staging cluster that runs at production scale overnight and on weekends for no reason. The three copies of a data set that exist because a migration was never finished. The caching layer that became redundant when somebody fixed the query. None of these get cheaper; they get deleted.

Elimination is also where the resistance is, because deleting something requires knowing it is unused, and nobody wants to be the person who deleted the thing that turned out to matter. That is an argument for good attribution and clear ownership, not for keeping everything.

A right-sized system drifts out of true. Traffic grows, features change the access pattern, someone adds an index, a dependency gets faster. None of this is anyone’s fault and all of it means the sizing decision made six months ago is now approximately wrong.

So cost optimization is a cadence, not a milestone. Systems that were optimized once and then left alone reliably show the same pattern: a sharp improvement, then eighteen months of steady erosion back to roughly where they started, followed by another cost reduction project.

The cadence does not have to be frequent. It has to exist, have an owner, and compare actual spend against a model that states what was expected, which is what COST 1 and COST 9 are for. Without the model, the review has nothing to measure against and becomes a search for obvious waste, which is the retroactive project again.