SUS 7. How do you eliminate work rather than adding capacity for it?
Zuletzt aktualisiert am
PERF 6 makes this argument for speed and this question makes it for consumption. They share the
lever almost entirely, which is unusual and worth exploiting: an efficiency improvement made for
one reason counts for both without further justification.
The distinction that matters is where in the stack the improvement happens. A fix at the top removes work permanently, needs no maintenance and never regresses. Adding capacity underneath the same work consumes resources for as long as the system exists, and hides the inefficiency so that the next enlargement helps less than the last.
Best practices
Section titled “Best practices”SUS 7.1Remove the work before making it efficientSUS 7.2Fix at the highest layer where the fix is possibleSUS 7.3Prefer improvements that need no maintenanceSUS 7.4Count the improvement once and claim it for both pillars
SUS 7.1 Remove the work before making it efficient
Section titled “SUS 7.1 Remove the work before making it efficient”Risk if not established: Medium
Optimizing work that does not need to happen is a smaller version of the same waste. The question that comes first is whether the work is needed at all, and it is asked less often because nothing in a normal process asks it.
The recurring finds are the same as PERF 6.1: columns selected and never read, payloads
serialized and discarded, values recomputed that have not changed, results fetched that a caller
ignores, and validation performed three times because three layers each defend themselves.
Each is invisible as a problem, because the code doing it is correct and fast. What is visible is a system doing a great deal of work, all of it efficient, much of it pointless.
This overlaps with SUS 3.4 and with SUS 6, deliberately. Work nobody consumes, capacity nobody
uses and data nobody reads are one finding reached through three routes, and whichever route
surfaces it, removing it is the same action.
On STACKIT. No platform feature identifies unnecessary work. Distributed tracing through Observability makes it visible, because a trace shows the calls a request actually made rather than the ones anybody believes it makes.
What removal buys is a smaller variant under SUS 2, fewer nodes under SUS 4, and in the best
case a component that no longer needs to exist.
Tradeoffs. Almost none, which is what makes this the first move. Removing work reduces consumption, cost and complexity at once.
Verify. For the most expensive operation in your workload, which of the work it performs is consumed by anything? What proportion could be removed without any user noticing?
SUS 7.2 Fix at the highest layer where the fix is possible
Section titled “SUS 7.2 Fix at the highest layer where the fix is possible”Risk if not established: Medium
The same outcome can be reached at several layers, and the layers differ by orders of magnitude in what they cost and what they consume.
A query that stops reading data it never uses removes work permanently. An index makes the same query cheaper and adds a maintenance cost on every write. A larger performance class makes the same query fast and consumes more hardware continuously. A caching layer makes it appear fast and adds a component with its own consumption and its own failure modes.
All four are legitimate and they are not equivalent. The first is free after it is done; the last three are ongoing.
Work down that list rather than up it. The instinct under time pressure is to reach for the
capacity answer because it is fastest to apply, and that decision is permanent unless somebody
revisits it, which PERF 9.3 asks for and rarely happens.
Where the higher-layer fix is genuinely unavailable, the capacity answer is correct. Recording why is what allows a later review to distinguish a considered decision from a shortcut.
On STACKIT. The layers where the capacity answer is expensive are the ones to identify up
front. A
managed database performance class cannot be changed in place per PERF 3.2, so reaching for it
is a migration rather than a setting, which makes the higher-layer fix comparatively cheaper than
it first appears.
Tradeoffs. Operational Excellence. The higher-layer fix usually takes longer to implement and requires understanding the system rather than its configuration, which is why the capacity answer wins under deadline pressure.
Verify. For your last three capacity increases, was a higher-layer fix considered? What prevented it?
SUS 7.3 Prefer improvements that need no maintenance
Section titled “SUS 7.3 Prefer improvements that need no maintenance”Risk if not established: Medium
Efficiency improvements differ in whether they keep working without attention, and that difference compounds over the life of a system.
Removed work stays removed. A corrected algorithm keeps being correct. A cache needs invalidation, tuning and a decision about staleness. An autoscaler needs thresholds that somebody maintains. A compression setting needs revisiting when the data changes shape.
The maintained kind is not wrong, and it carries an obligation that the unmaintained kind does
not. An estate accumulating maintained optimizations accumulates a review burden, and the ones
that stop paying are rarely removed, which is PERF 9.3.
Prefer the durable kind where both are available. Where only the maintained kind will do, record why it exists, so the review that eventually asks whether it still pays has something to work from.
On STACKIT. Managed services are the version of this that applies at the infrastructure layer:
efficiency improvements in a managed service arrive without you doing anything, and its capacity
is shared across tenants rather than reserved per customer, which is SUS 4.1 operating at a
scale you do not control.
That is a genuine argument for the managed option beyond the labour comparison in COST 4.2, and
it is the one most likely to be overlooked because it is invisible from inside the workload.
Tradeoffs. Performance Efficiency. Some of the largest improvements are the maintained kind, particularly caching, so this best practice is a preference rather than a rule.
Verify. List the efficiency mechanisms in your workload. Which of them require ongoing attention, and when was each last reviewed?
SUS 7.4 Count the improvement once and claim it for both pillars
Section titled “SUS 7.4 Count the improvement once and claim it for both pillars”Risk if not established: Low
The consumption argument is hard to fund on its own: its benefit is diffuse, delayed and accrues to people who are not in the room, as the Sustainability tradeoffs describe.
Where an improvement also serves performance or cost, that does not matter: the work gets funded on the stronger argument and the outcome is identical.
The overlaps worth using: every removal under SUS 7.1 is PERF 6.1. Every idle resource under
SUS 6 is COST 5 and frequently SEC 1.4. Every unnecessary copy under SUS 5.3 is SEC 3.3.
Reserve the consumption-only argument for where it is the only one available, which is the
divergences named in SUS 1.4: committed capacity, cheap storage removing the pressure to delete,
and the point where cost stops caring.
On STACKIT. No platform feature applies. What helps is that the same measurements serve all
three: utilization from
Observability and
spend from the Cost
Dashboard are the inputs to
COST 3, PERF 9.1 and this pillar alike, which is the practical case for the joint review in
COST 9.4.
Tradeoffs. None. This is a framing decision about how the work gets justified rather than about what work is done.
Verify. Of the efficiency work done in the last year, how much was justified on cost or performance grounds? Was any of it justified on consumption grounds alone, and did it happen?
Related
Section titled “Related”PERF 6Reducing work, which shares this lever entirelySUS 3.4andSUS 6, which reach the same finding through different routesPERF 9.3Removing expired optimizations, the maintenance burden this tries to avoidCOST 9.4Joint review, where the shared measurements are examined together- Sustainability tradeoffs, which sets out where this pillar stands alone