---
id: REL08
pillar: reliability
title: REL 8. How do you back up data, and how do you know you can restore it?
description: An untested backup is not a backup. How to derive schedules from the RPO, place copies out of reach, and prove restores by performing them on a fixed cadence.
status: draft
services: [compute-engine, postgresql-flex, object-storage, kubernetes-engine]
sidebar:
  order: 17
  label: Backup & restore
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/rel-08-backup-and-restore/"
source_file: "docs/architecture/pillars/reliability/rel-08-backup-and-restore.mdx"
---

Backups fail quietly. The job that stopped running six weeks ago, the volume excluded when someone
renamed it, the snapshot that completes successfully and cannot be read. Managed backup services
increasingly do notify on failure, STACKIT's among them, which removes the worst version of this.
What no notification tells you is whether the backup is restorable, and that is the part
discovered during recovery, at the most expensive possible moment.

The distinguishing property of a good backup practice is not the schedule. It is whether anyone
has restored from it recently and timed how long it took.

## Best practices

- [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set) Derive the backup schedule from the RPO, per data set
- [`REL 8.2`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-82-hold-copies-where-a-single-failure-cannot-destroy-both) Hold copies where a single failure cannot destroy both
- [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it) Restore on a cadence into a clean environment, and time it
- [`REL 8.4`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-84-know-which-backups-the-platform-makes-and-which-are-yours) Know which backups the platform makes and which are yours
- [`REL 8.5`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-85-protect-backups-to-the-classification-of-their-contents) Protect backups to the classification of their contents

---

## REL 8.1 Derive the backup schedule from the RPO, per data set

**Risk if not established:** High

Backup frequency is not a preference. It follows arithmetic: the interval between backups is the
maximum data loss, so an RPO of fifteen minutes and a nightly backup are incompatible regardless
of how good the backup is.

Do this per data set rather than per system. A transactional database and a static asset store
frequently sit in the same workload with RPOs three orders of magnitude apart, and a single
schedule is wrong for one of them.

Retention is a separate decision from frequency and needs its own reasoning. Frequency covers
recovery from failure; retention covers recovery from a mistake discovered late, such as a
corruption that propagated for a week before anyone noticed. Regulatory retention is a third
requirement again, and [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) governs that one.

Continuous approaches change the arithmetic. Transaction log shipping and point-in-time recovery
reduce the loss window to seconds without a backup running every few seconds, at the cost of a
more involved restore procedure.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/basics/architecture-of-postgresql-flex/">PostgreSQL
Flex</LinkChip>
enables transaction logs by default, which is what makes point-in-time recovery possible across
the available snapshots. Daily backup timing is adjustable, and <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/how-tos/backup-and-clone-postgresql-flex/">backup and
clone</LinkChip>
covers the operations. The other managed databases have equivalent how-to pages in the <LinkChip href="https://docs.stackit.cloud/products/">product
documentation</LinkChip>.

> From the STACKIT docs: [Architecture › Backup](https://docs.stackit.cloud/products/databases/postgresql-flex/basics/architecture-of-postgresql-flex/#backup) (Source updated 19.03.2026, copied 05.10.2026)

You manage backups at the instance level, which includes all users and databases. Because the transaction logs are activated by default, you can use them to achieve point-in-time recovery across all available snapshots.

With **instance cloning**, you can replicate your existing PostgreSQL Flex instance to another instance within the same project for testing, staging, or data recovery workflows. Recovery to the existing instance is not possible yet.

For virtual machines, <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/">Server Backup
Management</LinkChip> provides
crash-consistent backups on configurable schedules with retention. Note the term crash-consistent:
it captures the disk as if power had been cut, which is fine for many workloads and not sufficient
for a database that needs application-consistent state. Where consistency matters, back up through
the database rather than under it.

Two Server Backup Management limits belong in the retention arithmetic rather than being
discovered by <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/how-tos/monitoring-server-backup/">an
alert</LinkChip>:
a maximum of 350 backups, and a backup storage quota. Frequency multiplied
by retention has to fit inside both, which is what decides whether a short interval and a long
retention are simultaneously possible.

PostgreSQL Flex retains backups for somewhere between 32 and 90 days. Read that as the window a
late-discovered corruption has to be found inside, because past it the good copy is gone. It is
generous compared with the seven or fourteen days teams often assume, and it is still shorter than
a retention obligation measured in years: where one exists, it is met by exporting rather than by
the service's own backups, and that export is a thing you build.

**Tradeoffs.** **Cost Optimization.** Backup storage accumulates and is rarely reviewed, which is
[`COST 6`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/). **Performance Efficiency.** Backup windows consume I/O, which is why they are scheduled
and why the schedule sometimes conflicts with a batch window.

**Verify.** For each data set, what is the RPO and what is the backup interval? Where the interval
exceeds the RPO, who accepted that gap?

---

## REL 8.2 Hold copies where a single failure cannot destroy both

**Risk if not established:** High

A backup stored next to the thing it protects shares the failure it was meant to survive. This is
obvious for a snapshot on the same volume and less obvious for a backup in the same zone, the same
project, or under the same credentials.

Think in terms of what one event can reach. Physical failure argues for a different zone or
region. Accidental deletion argues for a different project and separate permissions. Malicious
action argues for immutability, because an attacker with your credentials will delete backups
first.

The credential boundary is the one most often missed. If the identity running the workload can
also delete the backups, then a compromise or a scripting error takes both, and the backup
provided no independence at all.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/">Object Storage</LinkChip> is
the usual destination for copies that must outlive a zone, and <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/how-tos/manage-bucket-versioning/">bucket
versioning</LinkChip>
protects against overwrite and deletion within it.

For retention that must be provably unaltered,
<LinkChip href="https://docs.stackit.cloud/products/storage/archiving/">Archiving</LinkChip> provides audit-proof immutable
storage, which is the right tool where [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) requires tamper-evidence rather than merely a copy.

Separation of permissions comes from the resource hierarchy: holding backups in a different
project with distinct role assignments is what makes the credential boundary real, which is the
same structure [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) asks for.

For a second region, `eu01` and `eu02` are both available and both inside EU jurisdiction, so
cross-region copies do not force a sovereignty tradeoff here the way they can elsewhere. [`SOV 2`](/architecture/pillars/sovereignty/sov-02-placement-and-residency/)
still requires the placement to be recorded.

**Tradeoffs.** **Cost Optimization.** Multiple copies in multiple places, each with its own
retention. **Sovereignty & Compliance.** Every copy inherits the classification of its contents
and must satisfy [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/), including its location and retention.

**Verify.** For your most critical data set, where do the copies live, and can the identity that
runs the workload delete them? What single event would destroy both the primary and the backup?

---

## REL 8.3 Restore on a cadence into a clean environment, and time it

**Risk if not established:** High

This is the best practice that separates a backup practice from a backup configuration, and it is
the one most often skipped.

Restore into a clean environment rather than over the original. Restoring on top of a working
system usually succeeds because the pieces you forgot to back up are still there, which is exactly
the property you are trying to test. The clean environment is what exposes the missing
configuration, the undocumented dependency and the credential nobody captured.

Time it, and compare the result with the RTO from [`REL 1.2`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-12-set-availability-rto-and-rpo-per-critical-flow-rather-than-per-workload). The restore duration is not the copy
duration: it includes locating the right backup, provisioning the target, restoring, verifying,
and reconnecting the dependencies. Teams routinely discover that a four-hour RTO is backed by a
nine-hour procedure.

Verify the restored data rather than the exit code. A restore that completes and produces an empty
or truncated data set is a successful backup job and a failed recovery.

Set a cadence proportional to consequence: quarterly for the most critical data sets, and after
any material change to the schema, the platform version or the backup configuration.

**On STACKIT.** Restore procedures are documented per service, such as <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/how-tos/backup-and-clone-postgresql-flex/">backup and clone for
PostgreSQL
Flex</LinkChip>,
and the ability to clone an instance is useful precisely because it gives you a clean target
without disturbing production.

For the detection half, <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/how-tos/monitoring-server-backup/">monitoring server
backup</LinkChip>
sends alert emails to project owners, admins and members when a scheduled backup cannot run or a
quota is exceeded. That covers the failing-job case without you building anything. Treat it as a
signal to route rather than a control in itself: an
email to a group of people is read by whoever happens to look, and the alerting in [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) is
where it belongs if a missed backup matters.

One constraint belongs in the design early: for PostgreSQL Flex, restoring into an existing
instance is not currently supported, so recovery means creating a new instance and repointing
consumers. That is a normal shape for managed databases and it changes the recovery procedure, so
the repointing step belongs in the rehearsal and in the RTO rather than being discovered during
one.

For Kubernetes, cluster resources are a separate concern from the data in persistent volumes.
<LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/storage/backup-management/">Backup
management</LinkChip>
and a <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/backup-your-cluster-with-velero/">documented cluster backup
path</LinkChip>
cover the cluster side; a rehearsal that restores the data and not the cluster definition has
tested half of the recovery.

> From the STACKIT docs: [Backup management › Cluster data backup](https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/storage/backup-management/#cluster-data-backup) (Source updated 19.08.2026, copied 05.10.2026)

Kubernetes clusters can vary a lot in terms of workloads and data they contain. Therefore, we cannot provide a central backup solution for data that is used and/or produced by the applications deployed in your cluster. Precisely, this affects the following:

- Data inside persistent volumes
- Data stored on the worker nodes
- Any data inside your container that is not part of the container image

The last two bullet points are considered an anti-pattern, anyway. Whenever possible, you should build and use stateless containers. If stateful data is necessary for your application, persistent volumes should be used. Backing those up is the customer’s responsibility.

**Tradeoffs.** **Operational Excellence.** Real engineering time on a recurring basis, producing
nothing a customer sees. It is the clearest example of the compounding argument in the
[Operational Excellence tradeoffs](/architecture/pillars/operational-excellence/tradeoffs/).

**Verify.** When did you last restore this data set into a clean environment, how long did it take
end to end, and how did that compare with the RTO?

---

## REL 8.4 Know which backups the platform makes and which are yours

**Risk if not established:** High

The most damaging backup gap is the one nobody knew existed, and it usually comes from an
assumption that a managed service covers something it does not.

The division of labour differs by service and it is documented, so this is a reading exercise
rather than a research project. For each component, establish what the provider backs up, what it
retains, how long, and what remains yours.

Two categories fall through consistently. **Configuration** is not data: the cluster definition,
the network layout, the secrets and the pipeline configuration are all needed for recovery and are
frequently backed up by nobody. This is one of several returns on [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/), since infrastructure as
code makes configuration recoverable by construction. And **anything you built yourself** on
compute you manage is entirely yours, which is obvious when stated and easy to overlook when a
workload has grown gradually.

**On STACKIT.** The Compute Engine <LinkChip href="https://stackit.com/en/gtc/service-certificates">service
certificate</LinkChip> is explicit that backup and recovery
of Compute Engine are the customer's responsibility and not included in the service. That is the
usual division for infrastructure services, and STACKIT offers <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/">Server Backup
Management</LinkChip> as a
managed option to cover it. The point is that your RPO should name which of the two it relies on.

Managed databases sit on the other side of the line: PostgreSQL Flex takes backups and enables
transaction logs by default. Managed does not mean unlimited, though, so the retention question
from [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set) still applies, and long-term retention beyond the service default is yours to
arrange.

<LinkChip href="https://docs.stackit.cloud/products/storage/file-storage/getting-started/create-resource-pool-snapshots/">File
Storage</LinkChip>
supports resource pool snapshots on a schedule, which is a different mechanism again.

The general rule when reading a service certificate: what it does not mention, it does not cover.

**Tradeoffs.** None. This is an inventory exercise and its only cost is the time to do it once and
revisit it when the architecture changes.

**Verify.** For every stateful component in your workload, write down who backs it up, with what
retention. Which entries did you have to guess at, and which say "nobody"?

---

## REL 8.5 Protect backups to the classification of their contents

**Risk if not established:** High

A backup is a full copy of your data, usually in a place with less attention than production. It
carries the same classification, the same regulatory obligations and the same attractiveness to an
attacker.

Three properties to carry across: encryption with keys you control where the classification
requires it, access restricted to the identities that genuinely need it, and residency that
satisfies the same rules as the primary data set.

Deletion is part of protection. A backup retained past its required period is a liability rather
than an asset, and it is exactly what [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/) and data minimization under [`SEC 3`](/architecture/pillars/security/sec-03-data-classification/) both address.

**On STACKIT.** Encryption at rest and key ownership are covered by [`SEC 7`](/architecture/pillars/security/sec-07-encryption/) and [`SOV 4`](/architecture/pillars/sovereignty/sov-04-key-ownership/), and
<LinkChip href="https://docs.stackit.cloud/products/security/kms/">KMS</LinkChip> is where customer-managed keys live.

One consequence deserves stating plainly because it cuts against the rest of this question: if you
hold your own keys, losing a key destroys the backup as completely as losing the backup. Key
material becomes the most critical state in the system, and it needs the same rigour applied here
to backups. That is a real cost of [`SOV 4`](/architecture/pillars/sovereignty/sov-04-key-ownership/) and it is named in the
[Sovereignty tradeoffs](/architecture/pillars/sovereignty/tradeoffs/).

Residency follows from [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/): a backup in a location your classification does not permit is a
compliance finding regardless of how well it serves recovery.

**Tradeoffs.** **Operational Excellence.** Key lifecycle management for backups is a practice, not
a setting. **Cost Optimization.** Encrypted, replicated, long-retained backups accumulate cost
that is easy to cut and hard to justify cutting.

**Verify.** Who can read your backups, who can delete them, and are they encrypted with keys held
where your classification requires? If a key were lost today, which backups become unreadable?

---

## Related

- [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) Reliability targets, which supply the RPO and RTO this question serves
- [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/) Redundancy, which covers failure but not mistake or corruption
- [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/) Disaster recovery, where restores happen under time pressure
- [`SEC 7`](/architecture/pillars/security/sec-07-encryption/) Encryption and [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/) Telemetry and backup residency
- [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/) Data lifecycle, which asks why old backups still exist
