---
type: tradeoffs
pillar: cost-optimization
code: COST
title: "Cost Optimization: tradeoffs"
description: "What cost pressure quietly takes from reliability, security, operations and sovereignty: usually without anyone deciding it should."
status: draft
sidebar:
  label: "Tradeoffs"
  order: 2
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/tradeoffs/"
source_file: "docs/architecture/pillars/cost-optimization/tradeoffs.mdx"
---

Cost is the pillar with the loudest advocate. Every other quality in this framework is defended by
the people who understand it; cost is defended by a budget, a finance function, and a number that
appears every month whether or not anyone is paying attention.

That asymmetry cuts both ways. Cost pressure is the most common reason other pillars get quietly
weakened (retention shortened, redundancy removed, an environment consolidated) usually without
the tradeoff being stated as a tradeoff. And cost is also the pillar most often sacrificed by
default, through architectures that were never priced at all.

---

## Against Reliability

**Redundancy against spend.** Covered from the other side in the [Reliability
tradeoffs](/architecture/pillars/reliability/tradeoffs/); the short version is that redundancy means paying for
capacity that produces nothing until it is needed, and the cost does not scale linearly with the
benefit.

The specific failure mode belongs here rather than there: reliability gets removed under cost
pressure *incrementally and invisibly*. Backup retention is shortened. A standby is downsized. A
third replica is dropped. Each change is individually defensible and none of them are announced as
a reduction in reliability, so the workload's actual recovery capability drifts away from its
stated targets without anyone deciding it should.

**How to resolve it:** [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/), targets agreed with the business, is the defence. A change that
breaches a stated RTO is a decision someone has to make explicitly. Without stated targets, every
reduction looks like efficiency.

## Against Security

Security costs are easy to cut because their value is invisible until it is urgently not.

Log retention is the clearest example: cost is visible monthly and growing, value appears once,
during an investigation, and is enormous. It gets shortened in cost reviews more often than any
other security control. Segmentation is second, separate projects and networks per environment and
sensitivity class prevent consolidation, and consolidating them saves real money and removes real
containment.

**How to resolve it:** record the *reason* for security cost decisions next to the number.
Retention set to a period because a regulation or an investigation requirement demands it is
defensible in a cost review. Retention set to a period because it was the default is not.

## Against Performance Efficiency

Mostly allies. Efficient code needs less capacity; a fixed query removes a caching layer; better
data modelling reduces both latency and storage. Most performance optimizations are cost
optimizations and vice versa.

They diverge at the edges:

- **Headroom.** Capacity that absorbs spikes is idle most of the time. Cutting it saves money
  until the spike.
- **Caching and pre-computation** trade storage and memory cost for latency. Sometimes the trade
  is bad in one direction, sometimes the other, and it changes as traffic changes.
- **Reserved or committed capacity** is cheaper per unit and less elastic, which is a cost win
  that becomes a performance constraint the moment demand moves.

**How to resolve it:** [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/), stated performance targets, plays the same role [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) does for
reliability. Without them, "fast enough" is whatever the current cost pressure permits.

## Against Operational Excellence

Two directions, and they point opposite ways.

**Cost pressure against operations:** managed services cost more per unit than self-operated
equivalents, and the comparison usually omits the operational labour, the on-call burden, and the
expertise the self-operated option requires. A managed database that looks expensive next to a
virtual machine is often cheaper once someone counts the engineer-hours.

**Operations against cost:** automation, tooling, and dashboards cost real engineering time and
produce nothing a customer sees. They are the first thing cut when a team is under delivery
pressure, and the absence compounds, including in cost, since attribution and right-sizing both
depend on automation nobody funded.

**How to resolve it:** include operational labour in the cost model. [`COST 1`](/architecture/pillars/cost-optimization/cost-01-cost-model/) should account for
who runs the thing, not only what it consumes. Total cost of ownership is a cliché because it is
routinely omitted.

## Against Sovereignty & Compliance

Sovereignty controls cost money (see [Sovereignty tradeoffs](/architecture/pillars/sovereignty/tradeoffs/)) and
the cost is concentrated in places that look optional from a spreadsheet: key management
operations, confidential computing premiums, extended audit retention, foregone proprietary
services, exit testing that produces nothing.

The distinctive risk is that these look like technical choices rather than compliance
requirements, so they are cut by people who do not know what they are cutting.

**How to resolve it:** mark compliance-driven costs as such in the attribution model, with the
requirement that drives them. A line item labelled "regulatory retention, sector requirement"
survives a cost review. One labelled "log storage" does not.

## Against Sustainability

Largely aligned, both want less consumed capacity, and the divergences are worth knowing precisely
because the alignment is assumed.

- **Reserved and committed capacity** is financially efficient and can be physically wasteful: you
  are paying less for resources you may not use, but they are still allocated.
- **Cheap storage tiers** reduce cost per byte, which reduces the pressure to delete.
  Sustainability wants the data gone; cost optimization is satisfied once it is cheap.
- **Idle non-production environments** are a cost problem and an environmental one at similar
  magnitude: this is where the two pillars agree most strongly.

**How to resolve it:** where they agree, act once and count it for both. Where they diverge,
sustainability generally wants *elimination* where cost optimization is satisfied with
*reduction*. [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/) and [`COST 6`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/) are the same lever with different stopping points.

---

## Related

- <LinkChip href="/architecture/pillars/cost-optimization/principles/">Design principles</LinkChip>
- [Overview](/architecture/pillars/cost-optimization/): the questions this pillar asks
- <LinkChip href="/architecture/pillars/reliability/tradeoffs/">Reliability tradeoffs</LinkChip>
