Zum Inhalt springen
Beta

Cost Optimization

Zuletzt aktualisiert am

Is the money buying business value?

Cost optimization is not cost reduction. A workload can be made arbitrarily cheap by making it useless, and the cheapest architecture is almost never the right one. The question this pillar asks is narrower and harder: for every unit of spend, is it buying something the business wants?

That reframing matters because the two goals lead to different work. Cost reduction is a project, usually launched under pressure, that finds the obvious waste and then stops. Cost optimization is a continuous practice that keeps spend aligned with value as both drift, and they always drift, because requirements change, traffic patterns change, and yesterday’s correct instance size is today’s waste without anyone having done anything wrong.

  • Building a cost model before building the workload
  • Attribution: knowing which team, product, or flow caused a cost
  • Right-sizing as a continuous activity rather than an event
  • Choosing services and tiers that match the actual requirement
  • Environments: non-production is where the quiet waste accumulates
  • Data lifecycle costs, which grow whether you attend to them or not
  • Directing spend by business value rather than by resource
  • Visibility, anomaly detection, and a review cadence

Efficiency in the sense of doing more work per unit of resource belongs to Performance Efficiency, the two pillars agree far more often than they conflict, and a performance optimization is usually also a cost optimization. Reducing resource consumption for environmental reasons belongs to Sustainability, which mostly agrees with this pillar and diverges in specific places worth knowing about.

The cost of not being reliable, secure, or compliant is real and is not modelled here; it appears in the tradeoffs of the pillars that address those risks.

Cost you cannot attribute is cost you cannot manage. This is the question everything else depends on, and the one most often deferred because it requires organizational work rather than technical work.

An invoice that says the platform cost a certain amount last month is not actionable. An invoice that says which product, which environment, and which team caused each part of it turns a finance conversation into an engineering one, and puts the information in front of the people who can act on it, which is the only place it produces change.

The second idea: the cheapest resource is the one you do not run. Optimization instinctively reaches for cheaper tiers and better discounts, which are real but bounded. The larger wins are structurally different, deleting environments nobody uses, shutting down capacity outside working hours, removing a caching layer that a query fix made unnecessary. Elimination beats discounting, and it is where the first pass should look.

  1. Design principles
  2. Tradeoffs

Nine questions. Numbers follow the order the decisions are usually made in and do not indicate priority.


COST 1: How do you estimate what a design will cost before you build it?

Section titled “COST 1: How do you estimate what a design will cost before you build it?”

Estimate what the design will cost, per environment, before it is built, and state what drives the number, so that later divergence is diagnosable rather than merely surprising. A design nobody priced has an unexamined requirement.

→ Best practices

COST 2: How do you attribute cost to the team or product that causes it?

Section titled “COST 2: How do you attribute cost to the team or product that causes it?”

Structure the resource hierarchy and labelling so that every cost traces to a product, an environment, and a team. Attribution is a prerequisite for everything else in this pillar and is expensive to retrofit.

→ Best practices

COST 3: How do you keep provisioned capacity matched to measured demand?

Section titled “COST 3: How do you keep provisioned capacity matched to measured demand?”

Compare provisioned capacity against measured usage on a cadence, and act on the gap. Sizing decisions decay; a system nobody has re-examined is a system that is now approximately the wrong size.

→ Best practices

COST 4: How do you choose the service and tier that matches the requirement?

Section titled “COST 4: How do you choose the service and tier that matches the requirement?”

Select the cheapest option that meets the stated requirement rather than the one that meets the requirement you might have later. Re-examine when the requirement changes, tier decisions are rarely revisited even when their justification has disappeared.

→ Best practices

COST 5: How do you size non-production environments for their real purpose?

Section titled “COST 5: How do you size non-production environments for their real purpose?”

A test environment does not need production topology, production redundancy, or production uptime. Size each environment for what it is actually used for, and shut it down when it is not being used.

→ Best practices

Data accumulates by default and is almost never deleted by default. Define retention per data set, tier or archive data as it cools, and delete what has no further use, including old backups, snapshots, and logs.

→ Best practices

Direct cost scrutiny by business value, using the flow ranking from REL 2. Optimizing a resource list top-down by price finds the largest line items, which is not the same as finding the spend that buys the least.

→ Best practices

Put current and trending cost in front of the engineers whose decisions cause it, and alert on unexpected increases quickly enough to act. A monthly invoice is a report; a same-week anomaly alert is a control.

→ Best practices

COST 9: How do you review actual spend against the model on a fixed cadence?

Section titled “COST 9: How do you review actual spend against the model on a fixed cadence?”

Compare reality against COST 1’s model on a defined schedule, with a named owner, and act on the delta, either by fixing the workload or by correcting the model. Without a cadence, optimization happens under budget pressure, which is the most expensive time to do it.

→ Best practices