---
id: SUS05
pillar: sustainability
title: SUS 5. How do you delete data you no longer need?
description: Storing data is not an event. It is a commitment to keep hardware occupied, powered and replicated for as long as the data exists, plus every copy of it.
status: draft
services: [object-storage, observability]
sidebar:
  order: 14
  label: Data lifecycle
source_url: "https://framework.stackit.cloud/architecture/pillars/sustainability/sus-05-data-lifecycle/"
source_file: "docs/architecture/pillars/sustainability/sus-05-data-lifecycle.mdx"
---

[`COST 6`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/) addresses the same data and is satisfied once it sits on cheap storage. This question is
not, because a cheaper tier changes the price and not the hardware occupancy.

That divergence is the clearest one in this pillar. Cost optimization is satisfied by reduction;
this question is satisfied by elimination, and the gap between those two is where most retained
data sits.

## Best practices

- [`SUS 5.1`](/architecture/pillars/sustainability/sus-05-data-lifecycle/#sus-51-set-retention-with-a-stated-reason-and-delete-on-schedule) Set retention with a stated reason and delete on schedule
- [`SUS 5.2`](/architecture/pillars/sustainability/sus-05-data-lifecycle/#sus-52-delete-rather-than-tier-where-there-is-no-further-use) Delete rather than tier where there is no further use
- [`SUS 5.3`](/architecture/pillars/sustainability/sus-05-data-lifecycle/#sus-53-find-the-copies-which-are-usually-the-larger-share) Find the copies, which are usually the larger share
- [`SUS 5.4`](/architecture/pillars/sustainability/sus-05-data-lifecycle/#sus-54-collect-less-rather-than-storing-less) Collect less rather than storing less

---

## SUS 5.1 Set retention with a stated reason and delete on schedule

**Risk if not established:** Medium

The default is retention forever, and it is a behavioural default rather than a technical one.
Nobody is ever criticized for keeping data. Deleting it requires establishing that it is not
needed, which is work and carries a small chance of being wrong.

Invert the default. Deletion happens on a schedule and retention is what requires a reason, which
means every data set carries a period and a justification for it.

Three sources produce a legitimate period, and they differ in how negotiable they are. A
**regulatory floor** from [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) is not negotiable. An **operational period** from [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set)
covers recovery from a mistake discovered late. An **analytical period** is the softest and
usually the largest, and it is the one where the honest answer is frequently shorter than the
current setting.

Automate the deletion. A retention policy that requires somebody to act is a retention policy that
does not run, and the data accumulates while the policy exists on paper.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/reference/lifecycle-configuration/">Object Storage lifecycle
configuration</LinkChip>
expresses expiry as a property of the bucket rather than as a job somebody runs, which is what
makes the schedule actually happen.

Telemetry retention is bounded by the <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/">Observability service
plan</LinkChip>
rather than being unlimited, which means it cannot grow indefinitely by accident. The metrics
retention is the one worth examining, since it can be extended considerably and rarely gets
reduced once it has been.

**Tradeoffs.** **Reliability.** Shorter retention narrows the recovery window for a problem
discovered late. **Sovereignty & Compliance.** The regulatory floor overrides this question
entirely, and deleting below it is a finding rather than a saving.

**Verify.** For each data set, what is the retention, what is the reason, and does deletion happen
automatically or when somebody remembers?

---

## SUS 5.2 Delete rather than tier where there is no further use

**Risk if not established:** Medium

Tiering is where this question and [`COST 6.2`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/#cost-62-tier-data-as-it-cools-rather-than-keeping-everything-warm) part company. Moving cold data to cheaper storage
reduces the price and leaves the data occupying hardware, which satisfies cost optimization
completely and this question not at all.

Tiering is correct where the data is genuinely still needed and rarely accessed. It is a
substitute for deletion where nobody has asked whether it is needed, and cheaper storage makes
that question easier to avoid.

So the sequence matters. Ask whether the data has further use first. If it does, tier it. If it
does not, delete it, and the cheaper tier is not an alternative to that decision.

The signal that tiering has become a substitute is data that has been on cold storage for years
without a single retrieval. At that point the archive is a decision nobody made, and the storage
is occupied for a purpose nobody can state.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/how-tos/manage-s3-lifecycle-configuration/">Lifecycle
configuration</LinkChip>
expresses both transitions and expiry, so the policy that moves data to a cheaper tier can also be
the policy that eventually removes it. Setting the transition without the expiry is the common
configuration and the one this best practice argues against.

<LinkChip href="https://docs.stackit.cloud/products/storage/archiving/">Archiving</LinkChip> is immutable by design, which
makes it the right destination where [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) requires tamper-evidence and the wrong one where the
data was merely inconvenient to delete. Data placed there cannot be removed early, so the decision
to archive is a commitment to keep it.

**Tradeoffs.** **Cost Optimization**, which is satisfied by tiering and indifferent to what
follows. This is one of the few best practices in the framework where the two pillars give
different answers to the same question.

**Verify.** For data on your coldest storage tier, when was any of it last retrieved? What would
happen if it did not exist?

---

## SUS 5.3 Find the copies, which are usually the larger share

**Risk if not established:** Medium

A primary data set has replicas, backups, snapshots, an analytics copy, a test environment copy,
and exports somebody made for a purpose that finished. Each occupies hardware, and the total
across them usually exceeds the primary.

Trace them from the data set rather than looking for stored data. Asking where this data goes has
a complete answer; asking where stored data is does not.

Two categories resist deletion, each in its own way. **Versioned storage** retains previous
versions, so deleting an object may not free anything, and the visible object count diverges from
the billed and occupied volume. **Immutable storage** cannot be deleted early, by design.

The copies that go unnoticed longest are the manual ones: a database clone created for a rehearsal
under [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it), an export for a migration, a snapshot taken before a risky change. Each was
deliberate and short-lived in intent, and nothing removes them afterwards.

**On STACKIT.** Bucket versioning under [`REL 8.2`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-82-hold-copies-where-a-single-failure-cannot-destroy-both) is the mechanism where deletion does not reclaim
space, which is correct behaviour and the first thing to check when a retention policy appears not
to be working.

Database clones are the manual copy most often forgotten, because the clone operation exists
precisely to make them easy and nothing in the workflow suggests removing them. [`SEC 3.3`](/architecture/pillars/security/sec-03-data-classification/#sec-33-find-the-copies-telemetry-backups-caches-test-data-and-exports) finds
the same list for a different reason, and one sweep serves both.

**Tradeoffs.** **Reliability.** Backups and replicas are copies whose existence is required by
[`REL 8`](/architecture/pillars/reliability/rel-08-backup-and-restore/), so this best practice targets the unmanaged copies rather than the deliberate ones.

**Verify.** For your largest data set, how many copies exist and what is the total volume across
them? Which of those copies has an owner who still needs it?

---

## SUS 5.4 Collect less rather than storing less

**Risk if not established:** Medium

Data never collected occupies nothing, needs no retention policy, and cannot be forgotten. It is
the only measure in this question with no ongoing cost, and it is available less often than the
others but is larger when it is.

Three questions per data set. Is it collected because it is needed, or because it was available?
Is the full fidelity needed, or would an aggregate, a sample or a hash serve? Is every field
needed, or is the record being stored whole because that was easier?

Telemetry is where this pays most, because its volume scales with the system rather than with
usage and because sampling is well understood. Full-fidelity traces of every request are rarely
necessary, and [`OPS 7.3`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-73-instrument-for-the-questions-you-cannot-predict) already argues for cardinality discipline for a different reason.

This overlaps with data minimization under [`SEC 3.4`](/architecture/pillars/security/sec-03-data-classification/#sec-34-minimize-what-you-hold), and the overlap is useful: the same
reduction lowers consumption, lowers cost, and reduces what a breach would expose. Three arguments
for one change is unusually easy to fund.

**On STACKIT.** Telemetry volume is bounded by the <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/">Observability service
plan</LinkChip>,
which sets ingestion and storage limits per plan. That turns collection volume into a plan
decision rather than an unbounded growth, and choosing a smaller plan is a forcing function for
sampling that works better than an intention to sample.

**Tradeoffs.** **Operational Excellence.** Sampled telemetry means the specific request you want
during an incident may not have been captured, which is a real cost and the reason adaptive
sampling, keeping errors and slow requests in full, is the usual compromise.

**Verify.** For your highest-volume data source, what proportion of what you collect has ever been
read? What would be lost by sampling it?

---

## Related

- [`COST 6`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/) Data lifecycle, which stops at a cheaper tier rather than at deletion
- [`SEC 3.3`](/architecture/pillars/security/sec-03-data-classification/#sec-33-find-the-copies-telemetry-backups-caches-test-data-and-exports) Copies and [`SEC 3.4`](/architecture/pillars/security/sec-03-data-classification/#sec-34-minimize-what-you-hold) Minimization, which reach the same conclusions
- [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set) Backup retention, which sets a floor this question cannot go below
- [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) Auditability, which sets the regulatory floor
- [`OPS 7.4`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-74-set-retention-by-the-value-of-each-signal-rather-than-uniformly) Telemetry retention, one of the largest data sets in most estates
