Sustainability: tradeoffs
Last updated on
Resource consumption is the hardest of the seven qualities to fund, for a structural reason: its benefit is diffuse, delayed, and accrues to people who are not in the room. Every other quality can point at a consequence someone in the organization will personally experience, an outage, a breach, an invoice, a complaint. This one cannot.
Which means it wins where it is aligned with something else and loses where it is not. That determines where the effort should go: take every win that comes free with cost or performance, and reserve the genuinely contested arguments for the cases where the divergence is large.
Against Reliability
Section titled “Against Reliability”Redundancy against consumption, and it does not resolve in sustainability’s favour.
Redundancy is deliberately reserved idle capacity. A standby that never serves a request consumes resources continuously to protect against an event that may not occur in the system’s lifetime. Multi-zone replication duplicates storage and generates inter-zone traffic. From this pillar’s perspective, this is a substantial ongoing footprint for no delivered work.
It is also correct, and REL 4 is not waived for environmental reasons. A workload that fails and
has to be rebuilt, re-run, and reconciled generally consumes more than the redundancy would have.
The conflict also runs the other way. Consolidation concentrates blast radius: one host or one
cluster failing affects more workloads, and fewer, larger nodes mean a larger blast radius per
node failure. SUS 4.1 and SUS 4.3 both name that cost, and Reliability is the pillar that
pays it.
How to resolve it: prefer active-active over active-passive wherever the workload allows, so
that redundant capacity does useful work instead of waiting. Where standby is unavoidable, keep it
minimal and scale it on failover rather than mirroring production continuously. And note that
REL 1’s business-agreed targets serve this pillar too: redundancy sized to a stated RTO is
justified; redundancy sized to general anxiety is waste in both directions.
Against Cost Optimization
Section titled “Against Cost Optimization”Mostly the closest ally in the framework, and the alignment is worth examining rather than assuming, since it is the argument most often used to skip consumption work entirely.
They agree strongly on: idle environments, unused capacity, oversized instances, redundant data copies, work that nobody consumes. These are the majority of available wins. They are unambiguous, and they need no environmental justification to be worth doing.
They diverge on:
- Committed capacity: cheaper per unit, fully allocated. Cost improves, consumption does not.
- Cheap storage tiers: reduce cost per byte and remove the pressure to delete.
COST 6is satisfied;SUS 5is not. - Rounding up under uncertainty: rational cost behaviour that produces permanent slack.
- Stopping conditions: cost optimization ends when spend is acceptable; consumption keeps mattering after that.
How to resolve it: take the aligned wins on the cost argument, since they are easier to fund that way and the outcome is identical. Reserve the sustainability-specific case for the divergences above, where it is the only argument available.
Against Performance Efficiency
Section titled “Against Performance Efficiency”A genuinely two-sided relationship. See the Performance Efficiency tradeoffs for the same ground.
Aligned: efficiency is the shared core. SUS 7 and PERF 6 are the same discipline: work
not done consumes nothing and takes no time.
Divergent: performance is frequently bought with resources. Headroom for spikes is idle capacity. Pre-computation trades storage for latency. Keeping instances warm to avoid cold starts consumes continuously to serve occasionally. Read replicas duplicate data to spread load.
How to resolve it: ask which stated target requires the resource. Latency headroom on a flow
with an agreed PERF 1 target is justified. The same headroom on a batch job nobody is waiting
for is waste wearing the appearance of engineering, and it is common, because headroom is rarely
revisited once provisioned.
Against Operational Excellence
Section titled “Against Operational Excellence”Two directions.
Against: telemetry has a footprint. Observability wants everything collected, correlated, and retained, and that data occupies storage for its whole retention period. Multi-version deployments briefly run two copies. Test environments and load-testing infrastructure consume resources producing nothing user-facing.
For, and more significantly: almost everything in this pillar depends on automation that
Operational Excellence builds. Automated shutdown of idle environments, lifecycle policies that
actually execute, right-sizing driven by measured usage, discovering what is unused at all, none
of it happens manually at scale. OPS 3 and OPS 10 are prerequisites for most of SUS 6.
How to resolve it: treat the dependency as the primary relationship. The telemetry footprint is small; the automation is what makes this pillar executable.
Against Sovereignty & Compliance
Section titled “Against Sovereignty & Compliance”Sovereignty wins these outright, and the conflicts are minor.
Extended audit retention keeps data alive for years: SOV 7 requires it and SUS 5 cannot
override it. Confidential computing consumes more energy per unit of work. Placement constraints
remove the option of running workloads where energy is cleanest, which is a real lever on other
platforms and largely theoretical within eu01 and eu02.
How to resolve it: regulatory retention is a floor, not a target. It constrains what may be
deleted, not what must be kept beyond it, and data kept past its required retention because nobody
looked is exactly what SUS 5 addresses.
Against Security
Section titled “Against Security”One structural conflict, and it is the subject of SUS 4.4: density pulls against isolation.
Every boundary that SEC 2.1 requires prevents consolidation across it, and the best practice
exists to make that a decision rather than a default. A workload isolated because its
classification requires it is justified capacity; one isolated by history is a consolidation
candidate nobody examined.
The rest is minor. Log retention and encryption overhead have a footprint; neither is large enough to inform a security decision.
One alignment worth noting: data that has been deleted cannot be breached. SUS 5 and data
minimization under SEC 3 point the same way.
Related
Section titled “Related”- Design principles
- Overview: the questions this pillar asks
- Cost Optimization tradeoffs