Skip to content
Beta

Sustainability: design principles

Last updated on

1. You cannot improve what you cannot measure, and measurement here is hard

Section titled “1. You cannot improve what you cannot measure, and measurement here is hard”

Every pillar says some version of this. It carries more weight here because the measurement problem is genuinely unsolved rather than merely neglected.

Cost has an invoice. Latency has a timer. Environmental footprint has proxies (allocated capacity, utilization, storage volume, data transfer) that stand in for a physical reality determined by facts you cannot see from inside a virtual machine: how the hardware is shared, how the facility is powered, what the marginal energy source was at the moment your job ran.

The honest response is not to give up on measurement, and not to adopt a number because it is available. It is to be explicit about what a metric represents and what it omits. A workload optimized against a proxy that does not track physical consumption has produced a reporting outcome. Sometimes that is what was asked for; it should at least be named accurately.

Start with what you can defend: allocated versus used capacity, storage volume and its growth rate, how much runs when nobody needs it. These are measurable, they correlate with consumption, and acting on them is unambiguously correct.

2. The greenest workload is the one that does not run

Section titled “2. The greenest workload is the one that does not run”

Efficiency asks how to do this work with less. This principle asks whether the work needs doing.

The second question has larger answers and is almost never asked, because nothing in a normal development process asks it. Systems accumulate: environments outlive projects, reports outlive their audience, pipelines keep processing data nobody reads, and scheduled jobs run for years after the feature they supported was removed. Each was justified once. Nothing revisits.

This is the highest-leverage question in the pillar and the one that requires the least engineering. It requires someone with authority to find out what is still needed and permission to switch off what is not, which is more an organizational problem than a technical one.

A server at 15% utilization does not consume 15% of the resources. Idle infrastructure draws substantial power, occupies physical space, and represents manufactured hardware whose footprint was paid at production time regardless of how much work it subsequently does.

Which makes utilization density the dominant architectural lever: consolidating workloads onto fewer, better-used resources reduces consumption far more than making an individual workload marginally more efficient.

This is where sustainability and cost optimization agree most strongly, and where sustainability and reliability disagree most sharply, redundancy is deliberately reserved idle capacity, and REL 4 will not be waived for this pillar. The resolution is to prefer active-active designs, so that redundant capacity does useful work rather than waiting.

4. Data has an ongoing footprint, not a one-time one

Section titled “4. Data has an ongoing footprint, not a one-time one”

Storing data is not an event. It is a commitment to keep hardware occupied, powered, cooled, and replicated for as long as the data exists, plus the copies: backups, snapshots, replicas, the copy in the analytics store, the copy someone extracted for a migration in 2023.

The default is retention, and the default is stronger here than anywhere else in this framework. Nobody is ever blamed for keeping data. Deleting it requires establishing that it is not needed, which is work, and carries a small risk of being wrong, which is career-relevant. So it accumulates, and cheap storage tiers make accumulating comfortable.

The discipline is to make deletion the scheduled default and retention the thing that requires a reason. That reason often exists (regulation, investigation, genuine analytical value) and it should be recorded per data set, so that what remains is what someone justified rather than what nobody got round to removing.

5. Sustainability and cost align: until they do not

Section titled “5. Sustainability and cost align: until they do not”

Most of the time, doing this pillar well saves money, and that alignment is the reason sustainability work gets funded at all. Use it.

But relying on it entirely means stopping exactly where cost stops. Committed capacity is cheaper and equally allocated. Cheap storage removes the incentive to delete. Rounding up instance sizes is rational under cost uncertainty. And cost optimization ends when the number is acceptable, which is not the same condition as consumption being justified.

The practical guidance is to know which case you are in. Where the pillars agree (idle environments, unused capacity, redundant data) act once and count it twice; those are the easy wins and there are more of them than most teams expect. Where they diverge, the consumption case has to be made on its own terms, and those are the places where it is the only case available.