COST 6. How do you manage the cost of data across its whole lifecycle?
Zuletzt aktualisiert am
Storage is the cost that grows without anyone deciding it should. Compute is provisioned by someone and appears in a review; data arrives continuously, is retained by default, and the bill rises by a small amount every month until somebody notices the total.
The asymmetry is behavioural rather than technical. Nobody is ever criticized for keeping data. Deleting it requires establishing that it is not needed, which is work, and carries a small chance of being wrong, which is career-relevant.
Best practices
Section titled “Best practices”COST 6.1Set a retention period per data set, with a stated reasonCOST 6.2Tier data as it cools rather than keeping everything warmCOST 6.3Delete what has no further use, including the copiesCOST 6.4Watch what grows without anyone deciding
COST 6.1 Set a retention period per data set, with a stated reason
Section titled “COST 6.1 Set a retention period per data set, with a stated reason”Risk if not established: Medium
Retention without a reason is retention forever, because there is never a moment when deleting becomes obviously correct.
Four sources produce a defensible period, and they are frequently confused. Regulatory
retention is a floor set by an obligation, which SOV 7 establishes and which cost cannot argue
with. Operational retention is how far back you might need to recover, which REL 8.1 sets.
Forensic retention is how far back a security investigation must be able to see, which
SEC 11.1 sets against attacker dwell time; its value appears once, during an investigation,
which is exactly why it is the retention most often shortened in a cost review. Analytical
retention is how much history has value, which is usually the softest and the largest.
Record the reason next to the number. A retention set because a regulation requires it survives a cost review; one set because it was the default does not, and the second is what most estates are carrying.
Set it per data set rather than per system. A single policy applied to transaction records and to debug logs is wrong for one of them by a wide margin.
On STACKIT. Retention is configured per service, so this is several decisions rather than one. Object Storage supports lifecycle configuration , which is what turns a stated retention into something that happens without anyone acting. Telemetry retention is set by the Observability service plan , where metrics and logs already default to different periods.
Backups have their own arithmetic, and the Server Backup Management limits from REL 8.1 bound
it: frequency multiplied by retention has to fit inside a maximum backup count and a storage
quota, so a long retention and a short interval are not simultaneously available beyond a point.
Tradeoffs. Reliability. Shorter retention narrows the window for recovering from a mistake
discovered late, which is exactly what REL 8.1 warns about when retention is cut in a cost
review. Security. Shortening security-log retention is the act this pillar’s own
tradeoffs page warns against most vividly, and SEC 11.1 is the counterparty.
Sovereignty & Compliance. The regulatory floor is not negotiable.
Verify. For each data set, what is the retention and what is the stated reason? How many are set to a default nobody chose?
COST 6.2 Tier data as it cools rather than keeping everything warm
Section titled “COST 6.2 Tier data as it cools rather than keeping everything warm”Risk if not established: Medium
Most data is read intensively for a short period and then rarely, if ever. Keeping it all on storage priced for immediate access means paying access performance for data nobody accesses.
Tiering moves data to cheaper storage as it ages, trading retrieval time and sometimes retrieval cost for a lower standing price. The trade is good when the data is genuinely cold and bad when something reads it regularly, which makes the access pattern the input rather than the age.
Automate the transition. A tiering policy that requires someone to run it is a tiering policy that runs once.
Be careful about what tiering does not solve. Cheaper storage removes the pressure to delete,
which is where this best practice and SUS 5 part company: cost is satisfied once the data is
cheap, and consumption is not.
On STACKIT. Object Storage lifecycle configuration expresses transitions and expiry as a policy on the bucket rather than as a job you run.
For data that has to be retained in a demonstrably unaltered form,
Archiving provides audit-proof immutable
storage with its own service
plans . That is the
right destination where SOV 7 requires tamper-evidence rather than merely a cheap copy, and its
immutability is a constraint as well as a feature: data placed there cannot be deleted early to
save money.
Tradeoffs. Performance Efficiency. Colder tiers have slower retrieval, which matters if
the access pattern was misjudged. Reliability. A restore from archived storage takes longer,
which belongs in the RTO under REL 1.2 rather than being discovered during a recovery.
Verify. For your largest data set, what proportion has been read in the last ninety days? What tier is the remainder on?
COST 6.3 Delete what has no further use, including the copies
Section titled “COST 6.3 Delete what has no further use, including the copies”Risk if not established: Medium
Deletion is the only measure in this question that removes the cost entirely rather than reducing it, and it is the one that happens least.
The copies are where the volume hides. A primary data set has replicas, backups, snapshots, a copy
in the analytics store, an export somebody made for a migration, and a clone in a test
environment. Each is billed, and SEC 3.3 finds the same list for a different reason.
Old backups deserve specific attention because they accumulate on a schedule and are governed by a retention nobody revisits. Snapshots are worse, because they are usually created manually for a specific reason and outlive it silently.
Make deletion the scheduled default and retention the thing that requires a reason, which inverts the behaviour described at the top of this question. That is an organizational change more than a technical one.
On STACKIT. Lifecycle policies on Object Storage handle expiry automatically, which removes the human step that otherwise does not happen.
Two things resist automated deletion by design and are worth knowing before relying on a policy.
Bucket versioning under REL 8.2 retains previous versions, so deleting an object does not
necessarily reclaim its storage. And
Archiving is immutable, which is the
point of it and means early deletion is not available.
Database clones created for a rehearsal under REL 8.3 are the case most often forgotten, because
they are created deliberately for a short purpose and nothing removes them afterwards.
Tradeoffs. Reliability. Deleting something that turns out to be needed is unrecoverable,
which is why the retention reason from COST 6.1 comes first. Sovereignty & Compliance.
Deletion below a regulatory floor is a finding rather than a saving.
Verify. How many snapshots and database clones exist in your estate, and when was each created? Which of them have an owner who still needs them?
COST 6.4 Watch what grows without anyone deciding
Section titled “COST 6.4 Watch what grows without anyone deciding”Risk if not established: Medium
Storage cost rises gradually, which means no single month’s increase is large enough to trigger attention. Over a year the total can double without anyone noticing a step change.
Track the growth rate rather than the total. A data set growing five percent a month doubles in about fifteen months, and knowing that now is worth more than knowing the current figure precisely.
The categories that grow silently: telemetry, which scales with the system rather than with usage; audit and activity records, which are append-only by nature; backups, whose total is frequency multiplied by retention; and anything with versioning enabled, where the visible object count and the billed volume diverge.
Set an expectation for each and alert when growth exceeds it, which is COST 8.2 applied to a
dimension that changes too slowly for an anomaly detector tuned to daily spend.
On STACKIT. The Cost Dashboard offers monthly, quarterly, half-yearly, yearly and user-defined ranges, and the longer ranges are the ones that make a growth trend visible. A month-on-month view of a slowly growing line looks flat; a yearly view does not.
Telemetry is the case where growth is bounded by configuration rather than by behaviour, since the Observability service plans set both retention and storage limits. That makes it predictable, and it also means a plan chosen for today’s volume becomes a constraint rather than a cost as the system grows.
Tradeoffs. Little. This is a monitoring practice whose cost is the attention it requires.
Verify. For your three largest storage costs, what is the monthly growth rate? At that rate, what will they cost in a year?
Related
Section titled “Related”SUS 5Data lifecycle, the same lever stopping at deletion rather than at a cheaper tierREL 8.1Backup retention, which sets a floor this question cannot go belowSOV 7Auditability, which sets the regulatory floorSEC 11.1Security signals, whose evidence window is set against dwell time rather than costSEC 3.3Copies, which finds the same duplicated dataOPS 7.4Telemetry retention, one of the fastest-growing data sets in most estates