Skip to content
Beta

Cost Optimization: tradeoffs

Last updated on

Cost is the pillar with the loudest advocate. Every other quality in this framework is defended by the people who understand it; cost is defended by a budget, a finance function, and a number that appears every month whether or not anyone is paying attention.

That asymmetry cuts both ways. Cost pressure is the most common reason other pillars get quietly weakened (retention shortened, redundancy removed, an environment consolidated) usually without the tradeoff being stated as a tradeoff. And cost is also the pillar most often sacrificed by default, through architectures that were never priced at all.


Redundancy against spend. Covered from the other side in the Reliability tradeoffs; the short version is that redundancy means paying for capacity that produces nothing until it is needed, and the cost does not scale linearly with the benefit.

The specific failure mode belongs here rather than there: reliability gets removed under cost pressure incrementally and invisibly. Backup retention is shortened. A standby is downsized. A third replica is dropped. Each change is individually defensible and none of them are announced as a reduction in reliability, so the workload’s actual recovery capability drifts away from its stated targets without anyone deciding it should.

How to resolve it: REL 1, targets agreed with the business, is the defence. A change that breaches a stated RTO is a decision someone has to make explicitly. Without stated targets, every reduction looks like efficiency.

Security costs are easy to cut because their value is invisible until it is urgently not.

Log retention is the clearest example: cost is visible monthly and growing, value appears once, during an investigation, and is enormous. It gets shortened in cost reviews more often than any other security control. Segmentation is second, separate projects and networks per environment and sensitivity class prevent consolidation, and consolidating them saves real money and removes real containment.

How to resolve it: record the reason for security cost decisions next to the number. Retention set to a period because a regulation or an investigation requirement demands it is defensible in a cost review. Retention set to a period because it was the default is not.

Mostly allies. Efficient code needs less capacity; a fixed query removes a caching layer; better data modelling reduces both latency and storage. Most performance optimizations are cost optimizations and vice versa.

They diverge at the edges:

  • Headroom. Capacity that absorbs spikes is idle most of the time. Cutting it saves money until the spike.
  • Caching and pre-computation trade storage and memory cost for latency. Sometimes the trade is bad in one direction, sometimes the other, and it changes as traffic changes.
  • Reserved or committed capacity is cheaper per unit and less elastic, which is a cost win that becomes a performance constraint the moment demand moves.

How to resolve it: PERF 1, stated performance targets, plays the same role REL 1 does for reliability. Without them, “fast enough” is whatever the current cost pressure permits.

Two directions, and they point opposite ways.

Cost pressure against operations: managed services cost more per unit than self-operated equivalents, and the comparison usually omits the operational labour, the on-call burden, and the expertise the self-operated option requires. A managed database that looks expensive next to a virtual machine is often cheaper once someone counts the engineer-hours.

Operations against cost: automation, tooling, and dashboards cost real engineering time and produce nothing a customer sees. They are the first thing cut when a team is under delivery pressure, and the absence compounds, including in cost, since attribution and right-sizing both depend on automation nobody funded.

How to resolve it: include operational labour in the cost model. COST 1 should account for who runs the thing, not only what it consumes. Total cost of ownership is a cliché because it is routinely omitted.

Sovereignty controls cost money (see Sovereignty tradeoffs) and the cost is concentrated in places that look optional from a spreadsheet: key management operations, confidential computing premiums, extended audit retention, foregone proprietary services, exit testing that produces nothing.

The distinctive risk is that these look like technical choices rather than compliance requirements, so they are cut by people who do not know what they are cutting.

How to resolve it: mark compliance-driven costs as such in the attribution model, with the requirement that drives them. A line item labelled “regulatory retention, sector requirement” survives a cost review. One labelled “log storage” does not.

Largely aligned, both want less consumed capacity, and the divergences are worth knowing precisely because the alignment is assumed.

  • Reserved and committed capacity is financially efficient and can be physically wasteful: you are paying less for resources you may not use, but they are still allocated.
  • Cheap storage tiers reduce cost per byte, which reduces the pressure to delete. Sustainability wants the data gone; cost optimization is satisfied once it is cheap.
  • Idle non-production environments are a cost problem and an environmental one at similar magnitude: this is where the two pillars agree most strongly.

How to resolve it: where they agree, act once and count it for both. Where they diverge, sustainability generally wants elimination where cost optimization is satisfied with reduction. SUS 5 and COST 6 are the same lever with different stopping points.