Sustainability
Last updated on
Is the resource consumption justified?
Cloud resources are billed, so consuming less is usually cheaper, and most consumption work can ride on the cost argument. The two diverge wherever price and physical consumption come apart, and they come apart routinely:
- Committed capacity is cheaper per unit and still fully allocated. Financially efficient, physically unchanged.
- Cheap storage tiers reduce cost per byte, which removes the pressure to delete. The data still occupies hardware.
- Rounding up is rational under cost uncertainty: the next size costs a little more and removes a risk, and produces consumption nobody uses.
- Cost optimization stops when spend is acceptable. Consumption keeps mattering after that.
The direction of travel matters too. Cost is satisfied by reduction; sustainability is satisfied
by elimination. COST 6 is content once data is on a cheap tier. SUS 5 wants to know why it
still exists.
What this pillar covers
Section titled “What this pillar covers”- Establishing what can actually be measured and reported
- Right-sizing for real demand rather than for comfort
- Demand shaping, moving, batching, and deferring work
- Utilization density instead of reserved idle headroom
- Data lifecycle: tiering, archiving, and deletion
- Shutting down what nobody is using
- Efficiency at the top of the stack, where the leverage is largest
What it does not cover
Section titled “What it does not cover”The financial dimension belongs to Cost Optimization, and where the two
agree this pillar defers rather than repeating. Doing more work per unit of resource belongs to
Performance Efficiency, which this pillar borrows from constantly:
PERF 6 and SUS 7 are close relatives.
Provider-level concerns (data centre efficiency, energy sourcing, hardware lifecycle) are not architecture decisions and are not in scope here. They belong in provider selection.
The central idea
Section titled “The central idea”The greenest workload is the one that does not run. Every question here is a variation on it: capacity nobody uses, environments nobody opens, data nobody reads, computation nobody consumes.
That is a different search than efficiency. Efficiency asks how to do this work with less. Sustainability asks first whether the work needs doing, and the answer is more often “no” than teams expect, because nothing in a normal development process ever asks the question. Systems accumulate; almost nothing removes.
The second idea is a caution. You cannot improve what you cannot measure, and measurement here is genuinely harder than in the other pillars. Cost has an invoice. Latency has a timer. Resource consumption has proxies of varying quality, and a workload’s actual environmental footprint depends on facts about the infrastructure that are not visible from inside a virtual machine.
Which makes SUS 1 a real question rather than a formality: establish what you can honestly
measure before setting targets against it. A workload optimized against a metric that does not
correspond to physical reality has achieved a reporting outcome, not an environmental one, and
distinguishing the two is most of the discipline.
Where to start
Section titled “Where to start”Questions
Section titled “Questions”Seven questions. Numbers follow the order the decisions are usually made in and do not indicate priority.
Where a question would repeat one from Cost Optimization or Performance Efficiency, it is not restated here. The cross-reference is given instead. What remains is what this pillar asks that the others do not.
SUS 1: How do you measure the resource consumption of your workload?
Section titled “SUS 1: How do you measure the resource consumption of your workload?”Determine which consumption signals are actually available to you, what each represents, and what it omits. Set targets against metrics you can defend rather than against numbers that happen to be reportable.
SUS 2: How do you provision for real demand rather than for comfort?
Section titled “SUS 2: How do you provision for real demand rather than for comfort?”Provision against measured usage rather than against the size that removes the need to think about it. Rounding up is rational under cost uncertainty and produces allocated capacity that does no work.
Extends [COST 3](/architecture/pillars/cost-optimization/): cost stops when the spend is acceptable,
this does not.
SUS 3: How do you shape demand so that less capacity is needed?
Section titled “SUS 3: How do you shape demand so that less capacity is needed?”Establish which work has a genuine deadline and which merely runs on a schedule nobody chose deliberately. Batching and deferring non-urgent work lowers peak capacity requirements, which is what actually determines how much infrastructure exists.
SUS 4: How do you increase utilization density?
Section titled “SUS 4: How do you increase utilization density?”Consolidate workloads onto fewer, better-utilized resources. Where redundancy requires spare capacity, prefer active-active designs so that it does useful work rather than waiting for a failure that may not come.
SUS 5: How do you delete data you no longer need?
Section titled “SUS 5: How do you delete data you no longer need?”Define a retention period per data set with a stated reason, and delete on schedule, including backups, snapshots, replicas, and the extracts nobody remembers making. Moving data to cheap storage reduces cost and changes nothing physical.
Extends [COST 6](/architecture/pillars/cost-optimization/), with a different stopping point.
SUS 6: How do you find and shut down what nobody uses?
Section titled “SUS 6: How do you find and shut down what nobody uses?”Find the environments, jobs, pipelines, and services that no longer serve anyone, and switch them off. Automate shutdown for anything with predictable idle periods rather than relying on someone remembering.
SUS 7: How do you eliminate work rather than adding capacity for it?
Section titled “SUS 7: How do you eliminate work rather than adding capacity for it?”An algorithmic fix, a corrected query, or a removed redundant call eliminates work permanently and needs no maintenance. Scaling hardware to accommodate inefficiency consumes resources for as long as the system exists.
Shares its mechanism with
[PERF 6](/architecture/pillars/performance-efficiency/perf-06-reduce-work/).
Related
Section titled “Related”- Design principles
- Tradeoffs
- Cost Optimization: where the two agree, act once and count it for both