SUS 6. How do you find and shut down what nobody uses?
Last updated on
The greenest workload is the one that does not run. Every other question in this pillar makes something more efficient; this one removes it, which is a different order of result.
It is also the question that needs the least engineering and the most permission. Finding what nobody uses is straightforward. Switching it off requires somebody willing to be the person who turned off the thing that turned out to matter.
Best practices
Section titled “Best practices”SUS 6.1Find what nobody uses, rather than waiting for someone to report itSUS 6.2Shut down on a schedule what has predictable idle periodsSUS 6.3Destroy rather than stop where the state can be reproducedSUS 6.4Make it recurring, because the accumulation is continuous
SUS 6.1 Find what nobody uses, rather than waiting for someone to report it
Section titled “SUS 6.1 Find what nobody uses, rather than waiting for someone to report it”Risk if not established: Medium
Unused resources are never reported, because the people who would notice are the people who stopped using them. Finding them is an active exercise.
Start from the inventory rather than from memory, which means enumerating what exists and
subtracting what has an owner and a purpose. The remainder is the finding, and it is the same
remainder that COST 2.4 and SEC 1.4 produce.
Look for the specific signatures rather than for idleness in general: environments from projects that ended, resources created by people who have left, volumes attached to nothing, load balancers with no backends, clusters with no workloads, and databases with no connections.
Distinguish idle from unused. A disaster recovery standby is idle by design and needed, per
SUS 2.3. A test environment nobody has opened in eight months is unused. Confusing the two
produces either a reliability incident or a permanent exemption for everything.
On STACKIT. The Resource Manager hierarchy is the enumeration, and every resource resides within a project, so nothing exists outside it to miss.
Utilization comes from Observability , and the Cost Dashboard is the practical entry point because a project with cost and no activity is visible without any instrumentation at all.
Tradeoffs. Reliability. Deleting something that turns out to be needed is the risk that
makes this question uncomfortable, which is why the ownership from COST 2.4 comes first: an
unclaimed resource can be removed with more confidence than an unlabelled one.
Verify. How many projects in your organization had cost but no measurable activity last month? Who owns each of them?
SUS 6.2 Shut down on a schedule what has predictable idle periods
Section titled “SUS 6.2 Shut down on a schedule what has predictable idle periods”Risk if not established: Medium
A non-production environment used during working hours is idle for roughly three quarters of the week. Shutting it down outside those hours removes most of its consumption without removing anything anybody uses.
Automate it rather than relying on somebody remembering, which is OPS 10.2. A schedule is a
small piece of automation whose return recurs daily.
The property that decides whether this works is what stopping actually releases. Compute is frequently reserved rather than consumed on demand, so a resource that appears to be off may still be holding capacity that nothing else can use.
Storage almost always continues regardless, so an environment shut down nightly still occupies its
volumes. That residual is the floor of what scheduling can achieve and the reason SUS 6.3 is the
stronger answer where it is available.
On STACKIT. This distinction is documented precisely and it matters more here than it does for cost. The Compute Engine service certificate defines shelving as stopping a machine with its resource reservation cancelled, and excludes shelving periods from the billed period.
The billing consequence is COST 5.2. The physical consequence is this best practice: a machine
that is merely stopped keeps its reservation, which means the capacity remains allocated to it and
unavailable to anything else. Only shelving releases it. An automated shutdown that stops without
shelving therefore reduces neither the bill nor the occupancy, and both look identical from
outside.
Tradeoffs. Operational Excellence. Start-up time before the environment is usable, and the automation is a thing to maintain. Reliability, mildly: an environment started on demand is one more thing that can fail to start.
Verify. For each non-production environment, how many hours per week is it running and how many is it used? Of the resources stopped outside those hours, which are shelved and which merely stopped?
SUS 6.3 Destroy rather than stop where the state can be reproduced
Section titled “SUS 6.3 Destroy rather than stop where the state can be reproduced”Risk if not established: Medium
A destroyed environment consumes nothing at all, which is a stronger result than any amount of scheduling around one that persists.
It requires the environment to be reproducible from definitions, which is OPS 3 and OPS 6.2.
Where that capability exists it pays repeatedly: here, in the recovery rehearsals under REL 8.3
and REL 9.3, and in the load testing under PERF 8.3.
The candidates are environments created for a purpose with a beginning and an end: a feature
branch, a load test, a migration rehearsal, a demonstration. The ones that resist are those
holding long-lived state that is expensive to reproduce, and those are better kept and scheduled
under SUS 6.2.
Destroy by default rather than on request. An environment created on demand and never destroyed is
a permanent environment nobody planned, and it will appear in the next sweep under SUS 6.1 as
something nobody claims.
On STACKIT. Creation and destruction as a pipeline step depend on the definitions being code, through the Terraform provider , Pulumi , the CLI or the API .
Creating a project per ephemeral environment makes destruction clean and its consumption separately visible, and the 2,500 project limit per organization is high enough that this is practical rather than extravagant.
Tradeoffs. Operational Excellence. On-demand creation is a capability to build and maintain and only pays where it is used often. Reliability. An environment that has to be recreated is unavailable while it is being recreated.
Verify. Which of your environments exist permanently, and what state do they hold that could not be reproduced from definitions?
SUS 6.4 Make it recurring, because the accumulation is continuous
Section titled “SUS 6.4 Make it recurring, because the accumulation is continuous”Risk if not established: Medium
Resources are created continuously and removed in occasional sweeps, which means the estate grows between sweeps regardless of how thorough each one is.
Set a cadence with an owner. Quarterly suits most organizations, and what matters is that it
happens rather than how often, in the same way COST 9.3 argues for the cost review.
Better than a cadence is a trigger at creation. An environment created with a stated end date, or
a resource created by automation that also removes it, does not need a sweep to find it. That is
SUS 6.3 in a different form and it scales where a manual review does not.
Record what each sweep finds and removes. It is one of the few places in this pillar where the result is directly countable, and that number is what justifies the cadence continuing to exist.
Expect resistance to be organizational rather than technical. Nobody objects to the principle and
somebody always objects to the specific resource, which is why the ownership work in COST 2.4
matters more than the finding itself.
On STACKIT. The Cost
API retrieves
per-project data programmatically, which makes a recurring report a scheduled job rather than a
manual assembly. Combined with utilization from
Observability , the
candidate list can be generated rather than compiled, which is what makes the cadence survivable
and is a good candidate for OPS 10.2.
Tradeoffs. Operational Excellence. A recurring review costs time, and the individual findings are small. The accumulation it prevents is not.
Verify. When was your last sweep for unused resources, what did it find, and what happened to the findings?
Related
Section titled “Related”COST 5Environments, the same lever with a financial motiveCOST 2.4Attribution completeness, which produces the same unclaimed listSEC 1.4Estate coverage, which finds it for a third reasonOPS 3Everything as code, without which destruction is not reversibleSUS 2.3Headroom against slack, which distinguishes idle by design from unused