Zum Inhalt springen
Beta

REL 8. How do you back up data, and how do you know you can restore it?

Zuletzt aktualisiert am

Backups fail quietly. The job that stopped running six weeks ago, the volume excluded when someone renamed it, the snapshot that completes successfully and cannot be read. Managed backup services increasingly do notify on failure, STACKIT’s among them, which removes the worst version of this. What no notification tells you is whether the backup is restorable, and that is the part discovered during recovery, at the most expensive possible moment.

The distinguishing property of a good backup practice is not the schedule. It is whether anyone has restored from it recently and timed how long it took.

  • REL 8.1 Derive the backup schedule from the RPO, per data set
  • REL 8.2 Hold copies where a single failure cannot destroy both
  • REL 8.3 Restore on a cadence into a clean environment, and time it
  • REL 8.4 Know which backups the platform makes and which are yours
  • REL 8.5 Protect backups to the classification of their contents

REL 8.1 Derive the backup schedule from the RPO, per data set

Section titled “REL 8.1 Derive the backup schedule from the RPO, per data set”

Risk if not established: High

Backup frequency is not a preference. It follows arithmetic: the interval between backups is the maximum data loss, so an RPO of fifteen minutes and a nightly backup are incompatible regardless of how good the backup is.

Do this per data set rather than per system. A transactional database and a static asset store frequently sit in the same workload with RPOs three orders of magnitude apart, and a single schedule is wrong for one of them.

Retention is a separate decision from frequency and needs its own reasoning. Frequency covers recovery from failure; retention covers recovery from a mistake discovered late, such as a corruption that propagated for a week before anyone noticed. Regulatory retention is a third requirement again, and SOV 7 governs that one.

Continuous approaches change the arithmetic. Transaction log shipping and point-in-time recovery reduce the loss window to seconds without a backup running every few seconds, at the cost of a more involved restore procedure.

On STACKIT. PostgreSQL Flex enables transaction logs by default, which is what makes point-in-time recovery possible across the available snapshots. Daily backup timing is adjustable, and backup and clone covers the operations. The other managed databases have equivalent how-to pages in the product documentation .

Aus der STACKIT-DokuArchitektur von PostgreSQL Flex › SicherungStand der Quelle 17.11.2025 · übernommen 05.10.2026

Sie verwalten Ihre Sicherungen auf Instanz-Ebene. Somit sind alle Ihre Benutzer und Datenbanken enthalten. Mit Hilfe der standardmäßig aktivierten Transaktionsprotokolle können Sie jeden Zeitpunkt für alle verfügbaren Snapshots wiederherstellen.

Mit dem Klonen von Instanzen können Sie Ihre bestehende PostgreSQL-Flex-Instanz auf eine andere Instanz innerhalb desselben Projekts für Test-, Staging- oder Datenwiederherstellungs-Workflows replizieren. Eine Wiederherstellung auf der bestehenden Instanz ist noch nicht möglich.

Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

For virtual machines, Server Backup Management provides crash-consistent backups on configurable schedules with retention. Note the term crash-consistent: it captures the disk as if power had been cut, which is fine for many workloads and not sufficient for a database that needs application-consistent state. Where consistency matters, back up through the database rather than under it.

Two Server Backup Management limits belong in the retention arithmetic rather than being discovered by an alert : a maximum of 350 backups, and a backup storage quota. Frequency multiplied by retention has to fit inside both, which is what decides whether a short interval and a long retention are simultaneously possible.

PostgreSQL Flex retains backups for somewhere between 32 and 90 days. Read that as the window a late-discovered corruption has to be found inside, because past it the good copy is gone. It is generous compared with the seven or fourteen days teams often assume, and it is still shorter than a retention obligation measured in years: where one exists, it is met by exporting rather than by the service’s own backups, and that export is a thing you build.

Tradeoffs. Cost Optimization. Backup storage accumulates and is rarely reviewed, which is COST 6. Performance Efficiency. Backup windows consume I/O, which is why they are scheduled and why the schedule sometimes conflicts with a batch window.

Verify. For each data set, what is the RPO and what is the backup interval? Where the interval exceeds the RPO, who accepted that gap?


REL 8.2 Hold copies where a single failure cannot destroy both

Section titled “REL 8.2 Hold copies where a single failure cannot destroy both”

Risk if not established: High

A backup stored next to the thing it protects shares the failure it was meant to survive. This is obvious for a snapshot on the same volume and less obvious for a backup in the same zone, the same project, or under the same credentials.

Think in terms of what one event can reach. Physical failure argues for a different zone or region. Accidental deletion argues for a different project and separate permissions. Malicious action argues for immutability, because an attacker with your credentials will delete backups first.

The credential boundary is the one most often missed. If the identity running the workload can also delete the backups, then a compromise or a scripting error takes both, and the backup provided no independence at all.

On STACKIT. Object Storage is the usual destination for copies that must outlive a zone, and bucket versioning protects against overwrite and deletion within it.

For retention that must be provably unaltered, Archiving provides audit-proof immutable storage, which is the right tool where SOV 7 requires tamper-evidence rather than merely a copy.

Separation of permissions comes from the resource hierarchy: holding backups in a different project with distinct role assignments is what makes the credential boundary real, which is the same structure SEC 2 asks for.

For a second region, eu01 and eu02 are both available and both inside EU jurisdiction, so cross-region copies do not force a sovereignty tradeoff here the way they can elsewhere. SOV 2 still requires the placement to be recorded.

Tradeoffs. Cost Optimization. Multiple copies in multiple places, each with its own retention. Sovereignty & Compliance. Every copy inherits the classification of its contents and must satisfy SOV 3, including its location and retention.

Verify. For your most critical data set, where do the copies live, and can the identity that runs the workload delete them? What single event would destroy both the primary and the backup?


REL 8.3 Restore on a cadence into a clean environment, and time it

Section titled “REL 8.3 Restore on a cadence into a clean environment, and time it”

Risk if not established: High

This is the best practice that separates a backup practice from a backup configuration, and it is the one most often skipped.

Restore into a clean environment rather than over the original. Restoring on top of a working system usually succeeds because the pieces you forgot to back up are still there, which is exactly the property you are trying to test. The clean environment is what exposes the missing configuration, the undocumented dependency and the credential nobody captured.

Time it, and compare the result with the RTO from REL 1.2. The restore duration is not the copy duration: it includes locating the right backup, provisioning the target, restoring, verifying, and reconnecting the dependencies. Teams routinely discover that a four-hour RTO is backed by a nine-hour procedure.

Verify the restored data rather than the exit code. A restore that completes and produces an empty or truncated data set is a successful backup job and a failed recovery.

Set a cadence proportional to consequence: quarterly for the most critical data sets, and after any material change to the schema, the platform version or the backup configuration.

On STACKIT. Restore procedures are documented per service, such as backup and clone for PostgreSQL Flex , and the ability to clone an instance is useful precisely because it gives you a clean target without disturbing production.

For the detection half, monitoring server backup sends alert emails to project owners, admins and members when a scheduled backup cannot run or a quota is exceeded. That covers the failing-job case without you building anything. Treat it as a signal to route rather than a control in itself: an email to a group of people is read by whoever happens to look, and the alerting in OPS 7 is where it belongs if a missed backup matters.

One constraint belongs in the design early: for PostgreSQL Flex, restoring into an existing instance is not currently supported, so recovery means creating a new instance and repointing consumers. That is a normal shape for managed databases and it changes the recovery procedure, so the repointing step belongs in the rehearsal and in the RTO rather than being discovered during one.

For Kubernetes, cluster resources are a separate concern from the data in persistent volumes. Backup management and a documented cluster backup path cover the cluster side; a rehearsal that restores the data and not the cluster definition has tested half of the recovery.

Aus der STACKIT-DokuBackup management › Cluster data backupStand der Quelle 19.08.2026 · übernommen 05.10.2026

Kubernetes clusters can vary a lot in terms of workloads and data they contain. Therefore, we cannot provide a central backup solution for data that is used and/or produced by the applications deployed in your cluster. Precisely, this affects the following:

  • Data inside persistent volumes
  • Data stored on the worker nodes
  • Any data inside your container that is not part of the container image

The last two bullet points are considered an anti-pattern, anyway. Whenever possible, you should build and use stateless containers. If stateful data is necessary for your application, persistent volumes should be used. Backing those up is the customer’s responsibility.

Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

Tradeoffs. Operational Excellence. Real engineering time on a recurring basis, producing nothing a customer sees. It is the clearest example of the compounding argument in the Operational Excellence tradeoffs.

Verify. When did you last restore this data set into a clean environment, how long did it take end to end, and how did that compare with the RTO?


REL 8.4 Know which backups the platform makes and which are yours

Section titled “REL 8.4 Know which backups the platform makes and which are yours”

Risk if not established: High

The most damaging backup gap is the one nobody knew existed, and it usually comes from an assumption that a managed service covers something it does not.

The division of labour differs by service and it is documented, so this is a reading exercise rather than a research project. For each component, establish what the provider backs up, what it retains, how long, and what remains yours.

Two categories fall through consistently. Configuration is not data: the cluster definition, the network layout, the secrets and the pipeline configuration are all needed for recovery and are frequently backed up by nobody. This is one of several returns on OPS 3, since infrastructure as code makes configuration recoverable by construction. And anything you built yourself on compute you manage is entirely yours, which is obvious when stated and easy to overlook when a workload has grown gradually.

On STACKIT. The Compute Engine service certificate is explicit that backup and recovery of Compute Engine are the customer’s responsibility and not included in the service. That is the usual division for infrastructure services, and STACKIT offers Server Backup Management as a managed option to cover it. The point is that your RPO should name which of the two it relies on.

Managed databases sit on the other side of the line: PostgreSQL Flex takes backups and enables transaction logs by default. Managed does not mean unlimited, though, so the retention question from REL 8.1 still applies, and long-term retention beyond the service default is yours to arrange.

File Storage supports resource pool snapshots on a schedule, which is a different mechanism again.

The general rule when reading a service certificate: what it does not mention, it does not cover.

Tradeoffs. None. This is an inventory exercise and its only cost is the time to do it once and revisit it when the architecture changes.

Verify. For every stateful component in your workload, write down who backs it up, with what retention. Which entries did you have to guess at, and which say “nobody”?


REL 8.5 Protect backups to the classification of their contents

Section titled “REL 8.5 Protect backups to the classification of their contents”

Risk if not established: High

A backup is a full copy of your data, usually in a place with less attention than production. It carries the same classification, the same regulatory obligations and the same attractiveness to an attacker.

Three properties to carry across: encryption with keys you control where the classification requires it, access restricted to the identities that genuinely need it, and residency that satisfies the same rules as the primary data set.

Deletion is part of protection. A backup retained past its required period is a liability rather than an asset, and it is exactly what SUS 5 and data minimization under SEC 3 both address.

On STACKIT. Encryption at rest and key ownership are covered by SEC 7 and SOV 4, and KMS is where customer-managed keys live.

One consequence deserves stating plainly because it cuts against the rest of this question: if you hold your own keys, losing a key destroys the backup as completely as losing the backup. Key material becomes the most critical state in the system, and it needs the same rigour applied here to backups. That is a real cost of SOV 4 and it is named in the Sovereignty tradeoffs.

Residency follows from SOV 3: a backup in a location your classification does not permit is a compliance finding regardless of how well it serves recovery.

Tradeoffs. Operational Excellence. Key lifecycle management for backups is a practice, not a setting. Cost Optimization. Encrypted, replicated, long-retained backups accumulate cost that is easy to cut and hard to justify cutting.

Verify. Who can read your backups, who can delete them, and are they encrypted with keys held where your classification requires? If a key were lost today, which backups become unreadable?


  • REL 1 Reliability targets, which supply the RPO and RTO this question serves
  • REL 4 Redundancy, which covers failure but not mistake or corruption
  • REL 9 Disaster recovery, where restores happen under time pressure
  • SEC 7 Encryption and SOV 3 Telemetry and backup residency
  • SUS 5 Data lifecycle, which asks why old backups still exist