---
id: COST06
pillar: cost-optimization
title: COST 6. How do you manage the cost of data across its whole lifecycle?
description: Data accumulates by default and is never deleted by default. Nobody is blamed for keeping something, which is why retention decisions do not happen alone.
status: draft
services: [object-storage, observability]
sidebar:
  order: 15
  label: Data lifecycle
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/cost-06-data-lifecycle/"
source_file: "docs/architecture/pillars/cost-optimization/cost-06-data-lifecycle.mdx"
---

Storage is the cost that grows without anyone deciding it should. Compute is provisioned by
someone and appears in a review; data arrives continuously, is retained by default, and the bill
rises by a small amount every month until somebody notices the total.

The asymmetry is behavioural rather than technical. Nobody is ever criticized for keeping data.
Deleting it requires establishing that it is not needed, which is work, and carries a small chance
of being wrong, which is career-relevant.

## Best practices

- [`COST 6.1`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/#cost-61-set-a-retention-period-per-data-set-with-a-stated-reason) Set a retention period per data set, with a stated reason
- [`COST 6.2`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/#cost-62-tier-data-as-it-cools-rather-than-keeping-everything-warm) Tier data as it cools rather than keeping everything warm
- [`COST 6.3`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/#cost-63-delete-what-has-no-further-use-including-the-copies) Delete what has no further use, including the copies
- [`COST 6.4`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/#cost-64-watch-what-grows-without-anyone-deciding) Watch what grows without anyone deciding

---

## COST 6.1 Set a retention period per data set, with a stated reason

**Risk if not established:** Medium

Retention without a reason is retention forever, because there is never a moment when deleting
becomes obviously correct.

Four sources produce a defensible period, and they are frequently confused. **Regulatory
retention** is a floor set by an obligation, which [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) establishes and which cost cannot argue
with. **Operational retention** is how far back you might need to recover, which [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set) sets.
**Forensic retention** is how far back a security investigation must be able to see, which
[`SEC 11.1`](/architecture/pillars/security/sec-11-detection-and-response/#sec-111-collect-security-relevant-signals-where-the-subject-cannot-alter-them) sets against attacker dwell time; its value appears once, during an investigation,
which is exactly why it is the retention most often shortened in a cost review. **Analytical
retention** is how much history has value, which is usually the softest and the largest.

Record the reason next to the number. A retention set because a regulation requires it survives a
cost review; one set because it was the default does not, and the second is what most estates are
carrying.

Set it per data set rather than per system. A single policy applied to transaction records and to
debug logs is wrong for one of them by a wide margin.

**On STACKIT.** Retention is configured per service, so this is several decisions rather than one.
Object Storage supports <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/reference/lifecycle-configuration/">lifecycle
configuration</LinkChip>,
which is what turns a stated retention into something that happens without anyone acting.
Telemetry retention is set by the <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/">Observability service
plan</LinkChip>,
where metrics and logs already default to different periods.

Backups have their own arithmetic, and the Server Backup Management limits from [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set) bound
it: frequency multiplied by retention has to fit inside a maximum backup count and a storage
quota, so a long retention and a short interval are not simultaneously available beyond a point.

**Tradeoffs.** **Reliability.** Shorter retention narrows the window for recovering from a mistake
discovered late, which is exactly what [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set) warns about when retention is cut in a cost
review. **Security.** Shortening security-log retention is the act this pillar's own
[tradeoffs page](/architecture/pillars/cost-optimization/tradeoffs/) warns against most vividly, and [`SEC 11.1`](/architecture/pillars/security/sec-11-detection-and-response/#sec-111-collect-security-relevant-signals-where-the-subject-cannot-alter-them) is the counterparty.
**Sovereignty & Compliance.** The regulatory floor is not negotiable.

**Verify.** For each data set, what is the retention and what is the stated reason? How many are
set to a default nobody chose?

---

## COST 6.2 Tier data as it cools rather than keeping everything warm

**Risk if not established:** Medium

Most data is read intensively for a short period and then rarely, if ever. Keeping it all on
storage priced for immediate access means paying access performance for data nobody accesses.

Tiering moves data to cheaper storage as it ages, trading retrieval time and sometimes retrieval
cost for a lower standing price. The trade is good when the data is genuinely cold and bad when
something reads it regularly, which makes the access pattern the input rather than the age.

Automate the transition. A tiering policy that requires someone to run it is a tiering policy that
runs once.

Be careful about what tiering does not solve. Cheaper storage removes the pressure to delete,
which is where this best practice and [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/) part company: cost is satisfied once the data is
cheap, and consumption is not.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/how-tos/manage-s3-lifecycle-configuration/">Object Storage lifecycle
configuration</LinkChip>
expresses transitions and expiry as a policy on the bucket rather than as a job you run.

For data that has to be retained in a demonstrably unaltered form,
<LinkChip href="https://docs.stackit.cloud/products/storage/archiving/">Archiving</LinkChip> provides audit-proof immutable
storage with its own <LinkChip href="https://docs.stackit.cloud/products/storage/archiving/basics/service-plans/">service
plans</LinkChip>. That is the
right destination where [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) requires tamper-evidence rather than merely a cheap copy, and its
immutability is a constraint as well as a feature: data placed there cannot be deleted early to
save money.

**Tradeoffs.** **Performance Efficiency.** Colder tiers have slower retrieval, which matters if
the access pattern was misjudged. **Reliability.** A restore from archived storage takes longer,
which belongs in the RTO under [`REL 1.2`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-12-set-availability-rto-and-rpo-per-critical-flow-rather-than-per-workload) rather than being discovered during a recovery.

**Verify.** For your largest data set, what proportion has been read in the last ninety days? What
tier is the remainder on?

---

## COST 6.3 Delete what has no further use, including the copies

**Risk if not established:** Medium

Deletion is the only measure in this question that removes the cost entirely rather than reducing
it, and it is the one that happens least.

The copies are where the volume hides. A primary data set has replicas, backups, snapshots, a copy
in the analytics store, an export somebody made for a migration, and a clone in a test
environment. Each is billed, and [`SEC 3.3`](/architecture/pillars/security/sec-03-data-classification/#sec-33-find-the-copies-telemetry-backups-caches-test-data-and-exports) finds the same list for a different reason.

Old backups deserve specific attention because they accumulate on a schedule and are governed by a
retention nobody revisits. Snapshots are worse, because they are usually created manually for a
specific reason and outlive it silently.

Make deletion the scheduled default and retention the thing that requires a reason, which inverts
the behaviour described at the top of this question. That is an organizational change more than a
technical one.

**On STACKIT.** Lifecycle policies on <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/reference/lifecycle-configuration/">Object
Storage</LinkChip>
handle expiry automatically, which removes the human step that otherwise does not happen.

Two things resist automated deletion by design and are worth knowing before relying on a policy.
Bucket versioning under [`REL 8.2`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-82-hold-copies-where-a-single-failure-cannot-destroy-both) retains previous versions, so deleting an object does not
necessarily reclaim its storage. And
<LinkChip href="https://docs.stackit.cloud/products/storage/archiving/">Archiving</LinkChip> is immutable, which is the
point of it and means early deletion is not available.

Database clones created for a rehearsal under [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it) are the case most often forgotten, because
they are created deliberately for a short purpose and nothing removes them afterwards.

**Tradeoffs.** **Reliability.** Deleting something that turns out to be needed is unrecoverable,
which is why the retention reason from [`COST 6.1`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/#cost-61-set-a-retention-period-per-data-set-with-a-stated-reason) comes first. **Sovereignty & Compliance.**
Deletion below a regulatory floor is a finding rather than a saving.

**Verify.** How many snapshots and database clones exist in your estate, and when was each
created? Which of them have an owner who still needs them?

---

## COST 6.4 Watch what grows without anyone deciding

**Risk if not established:** Medium

Storage cost rises gradually, which means no single month's increase is large enough to trigger
attention. Over a year the total can double without anyone noticing a step change.

Track the growth rate rather than the total. A data set growing five percent a month doubles in
about fifteen months, and knowing that now is worth more than knowing the current figure
precisely.

The categories that grow silently: telemetry, which scales with the system rather than with usage;
audit and activity records, which are append-only by nature; backups, whose total is frequency
multiplied by retention; and anything with versioning enabled, where the visible object count and
the billed volume diverge.

Set an expectation for each and alert when growth exceeds it, which is [`COST 8.2`](/architecture/pillars/cost-optimization/cost-08-cost-visibility/#cost-82-alert-on-anomalies-rather-than-waiting-for-an-invoice) applied to a
dimension that changes too slowly for an anomaly detector tuned to daily spend.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost
Dashboard</LinkChip> offers monthly,
quarterly, half-yearly, yearly and user-defined ranges, and the longer ranges are the ones that
make a growth trend visible. A month-on-month view of a slowly growing line looks flat; a yearly
view does not.

Telemetry is the case where growth is bounded by configuration rather than by behaviour, since the
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/">Observability service
plans</LinkChip>
set both retention and storage limits. That makes it predictable, and it also means a plan chosen
for today's volume becomes a constraint rather than a cost as the system grows.

**Tradeoffs.** Little. This is a monitoring practice whose cost is the attention it requires.

**Verify.** For your three largest storage costs, what is the monthly growth rate? At that rate,
what will they cost in a year?

---

## Related

- [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/) Data lifecycle, the same lever stopping at deletion rather than at a cheaper tier
- [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set) Backup retention, which sets a floor this question cannot go below
- [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) Auditability, which sets the regulatory floor
- [`SEC 11.1`](/architecture/pillars/security/sec-11-detection-and-response/#sec-111-collect-security-relevant-signals-where-the-subject-cannot-alter-them) Security signals, whose evidence window is set against dwell time rather than cost
- [`SEC 3.3`](/architecture/pillars/security/sec-03-data-classification/#sec-33-find-the-copies-telemetry-backups-caches-test-data-and-exports) Copies, which finds the same duplicated data
- [`OPS 7.4`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-74-set-retention-by-the-value-of-each-signal-rather-than-uniformly) Telemetry retention, one of the fastest-growing data sets in most estates
