---
id: SUS06
pillar: sustainability
title: SUS 6. How do you find and shut down what nobody uses?
description: Systems accumulate and almost nothing removes. The highest-return question in this pillar needs the least engineering and the most organizational permission.
status: draft
services: [compute-engine]
sidebar:
  order: 15
  label: Shutting down idle
source_url: "https://framework.stackit.cloud/architecture/pillars/sustainability/sus-06-shut-down-idle/"
source_file: "docs/architecture/pillars/sustainability/sus-06-shut-down-idle.mdx"
---

The greenest workload is the one that does not run. Every other question in this pillar makes
something more efficient; this one removes it, which is a different order of result.

It is also the question that needs the least engineering and the most permission. Finding what
nobody uses is straightforward. Switching it off requires somebody willing to be the person who
turned off the thing that turned out to matter.

## Best practices

- [`SUS 6.1`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-61-find-what-nobody-uses-rather-than-waiting-for-someone-to-report-it) Find what nobody uses, rather than waiting for someone to report it
- [`SUS 6.2`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-62-shut-down-on-a-schedule-what-has-predictable-idle-periods) Shut down on a schedule what has predictable idle periods
- [`SUS 6.3`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-63-destroy-rather-than-stop-where-the-state-can-be-reproduced) Destroy rather than stop where the state can be reproduced
- [`SUS 6.4`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-64-make-it-recurring-because-the-accumulation-is-continuous) Make it recurring, because the accumulation is continuous

---

## SUS 6.1 Find what nobody uses, rather than waiting for someone to report it

**Risk if not established:** Medium

Unused resources are never reported, because the people who would notice are the people who
stopped using them. Finding them is an active exercise.

Start from the inventory rather than from memory, which means enumerating what exists and
subtracting what has an owner and a purpose. The remainder is the finding, and it is the same
remainder that [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims) and [`SEC 1.4`](/architecture/pillars/security/sec-01-security-baseline/#sec-14-cover-the-whole-estate-including-the-parts-nobody-claims) produce.

Look for the specific signatures rather than for idleness in general: environments from projects
that ended, resources created by people who have left, volumes attached to nothing, load balancers
with no backends, clusters with no workloads, and databases with no connections.

Distinguish idle from unused. A disaster recovery standby is idle by design and needed, per
[`SUS 2.3`](/architecture/pillars/sustainability/sus-02-right-sizing/#sus-23-distinguish-headroom-from-slack). A test environment nobody has opened in eight months is unused. Confusing the two
produces either a reliability incident or a permanent exemption for everything.

**On STACKIT.** The
<LinkChip href="https://docs.stackit.cloud/platform/resource-manager/">Resource Manager</LinkChip> hierarchy is the
enumeration, and every resource resides within a project, so nothing exists outside it to miss.

Utilization comes from
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip>, and
the <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost Dashboard</LinkChip> is the
practical entry point because a project with cost and no activity is visible without any
instrumentation at all.

**Tradeoffs.** **Reliability.** Deleting something that turns out to be needed is the risk that
makes this question uncomfortable, which is why the ownership from [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims) comes first: an
unclaimed resource can be removed with more confidence than an unlabelled one.

**Verify.** How many projects in your organization had cost but no measurable activity last month?
Who owns each of them?

---

## SUS 6.2 Shut down on a schedule what has predictable idle periods

**Risk if not established:** Medium

A non-production environment used during working hours is idle for roughly three quarters of the
week. Shutting it down outside those hours removes most of its consumption without removing
anything anybody uses.

Automate it rather than relying on somebody remembering, which is [`OPS 10.2`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-102-automate-what-recurs-starting-with-the-riskiest-rather-than-the-most-frequent). A schedule is a
small piece of automation whose return recurs daily.

The property that decides whether this works is what stopping actually releases. Compute is
frequently reserved rather than consumed on demand, so a resource that appears to be off may still
be holding capacity that nothing else can use.

Storage almost always continues regardless, so an environment shut down nightly still occupies its
volumes. That residual is the floor of what scheduling can achieve and the reason [`SUS 6.3`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-63-destroy-rather-than-stop-where-the-state-can-be-reproduced) is the
stronger answer where it is available.

**On STACKIT.** This distinction is documented precisely and it matters more here than it does for
cost. The Compute Engine <LinkChip href="https://stackit.com/en/gtc/service-certificates">service certificate</LinkChip>
defines shelving as stopping a machine **with its resource reservation cancelled**, and excludes
shelving periods from the billed period.

The billing consequence is [`COST 5.2`](/architecture/pillars/cost-optimization/cost-05-environments/#cost-52-shut-down-what-is-not-being-used-and-know-what-stopping-actually-stops). The physical consequence is this best practice: a machine
that is merely stopped keeps its reservation, which means the capacity remains allocated to it and
unavailable to anything else. Only shelving releases it. An automated shutdown that stops without
shelving therefore reduces neither the bill nor the occupancy, and both look identical from
outside.

**Tradeoffs.** **Operational Excellence.** Start-up time before the environment is usable, and the
automation is a thing to maintain. **Reliability**, mildly: an environment started on demand is
one more thing that can fail to start.

**Verify.** For each non-production environment, how many hours per week is it running and how
many is it used? Of the resources stopped outside those hours, which are shelved and which merely
stopped?

---

## SUS 6.3 Destroy rather than stop where the state can be reproduced

**Risk if not established:** Medium

A destroyed environment consumes nothing at all, which is a stronger result than any amount of
scheduling around one that persists.

It requires the environment to be reproducible from definitions, which is [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) and [`OPS 6.2`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-62-build-every-environment-from-the-same-definitions).
Where that capability exists it pays repeatedly: here, in the recovery rehearsals under [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it)
and [`REL 9.3`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-93-rehearse-it-and-time-the-rehearsal-against-the-rto), and in the load testing under [`PERF 8.3`](/architecture/pillars/performance-efficiency/perf-08-load-testing/#perf-83-test-where-the-result-transfers-to-production).

The candidates are environments created for a purpose with a beginning and an end: a feature
branch, a load test, a migration rehearsal, a demonstration. The ones that resist are those
holding long-lived state that is expensive to reproduce, and those are better kept and scheduled
under [`SUS 6.2`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-62-shut-down-on-a-schedule-what-has-predictable-idle-periods).

Destroy by default rather than on request. An environment created on demand and never destroyed is
a permanent environment nobody planned, and it will appear in the next sweep under [`SUS 6.1`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-61-find-what-nobody-uses-rather-than-waiting-for-someone-to-report-it) as
something nobody claims.

**On STACKIT.** Creation and destruction as a pipeline step depend on the definitions being code,
through the <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-iac/stackit-terraform-provider/">Terraform
provider</LinkChip>,
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-iac/pulumi/">Pulumi</LinkChip>, the
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-cli/">CLI</LinkChip> or the
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-api/">API</LinkChip>.

Creating a project per ephemeral environment makes destruction clean and its consumption
separately visible, and the <LinkChip href="https://docs.stackit.cloud/platform/resource-manager/basics/limitations/">2,500 project
limit</LinkChip> per organization
is high enough that this is practical rather than extravagant.

**Tradeoffs.** **Operational Excellence.** On-demand creation is a capability to build and
maintain and only pays where it is used often. **Reliability.** An environment that has to be
recreated is unavailable while it is being recreated.

**Verify.** Which of your environments exist permanently, and what state do they hold that could
not be reproduced from definitions?

---

## SUS 6.4 Make it recurring, because the accumulation is continuous

**Risk if not established:** Medium

Resources are created continuously and removed in occasional sweeps, which means the estate grows
between sweeps regardless of how thorough each one is.

Set a cadence with an owner. Quarterly suits most organizations, and what matters is that it
happens rather than how often, in the same way [`COST 9.3`](/architecture/pillars/cost-optimization/cost-09-review-cadence/#cost-93-give-the-review-an-owner-and-a-cadence-proportional-to-the-spend) argues for the cost review.

Better than a cadence is a trigger at creation. An environment created with a stated end date, or
a resource created by automation that also removes it, does not need a sweep to find it. That is
[`SUS 6.3`](/architecture/pillars/sustainability/sus-06-shut-down-idle/#sus-63-destroy-rather-than-stop-where-the-state-can-be-reproduced) in a different form and it scales where a manual review does not.

Record what each sweep finds and removes. It is one of the few places in this pillar where the
result is directly countable, and that number is what justifies the cadence continuing to exist.

Expect resistance to be organizational rather than technical. Nobody objects to the principle and
somebody always objects to the specific resource, which is why the ownership work in [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims)
matters more than the finding itself.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/how-tos/retrieve-cost-data/">Cost
API</LinkChip> retrieves
per-project data programmatically, which makes a recurring report a scheduled job rather than a
manual assembly. Combined with utilization from
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip>, the
candidate list can be generated rather than compiled, which is what makes the cadence survivable
and is a good candidate for [`OPS 10.2`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-102-automate-what-recurs-starting-with-the-riskiest-rather-than-the-most-frequent).

**Tradeoffs.** **Operational Excellence.** A recurring review costs time, and the individual
findings are small. The accumulation it prevents is not.

**Verify.** When was your last sweep for unused resources, what did it find, and what happened to
the findings?

---

## Related

- [`COST 5`](/architecture/pillars/cost-optimization/cost-05-environments/) Environments, the same lever with a financial motive
- [`COST 2.4`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/#cost-24-make-attribution-complete-including-what-nobody-claims) Attribution completeness, which produces the same unclaimed list
- [`SEC 1.4`](/architecture/pillars/security/sec-01-security-baseline/#sec-14-cover-the-whole-estate-including-the-parts-nobody-claims) Estate coverage, which finds it for a third reason
- [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) Everything as code, without which destruction is not reversible
- [`SUS 2.3`](/architecture/pillars/sustainability/sus-02-right-sizing/#sus-23-distinguish-headroom-from-slack) Headroom against slack, which distinguishes idle by design from unused
