Zum Inhalt springen
Beta

REL 4. How do you eliminate single points of failure?

Zuletzt aktualisiert am

This is where the targets from REL 1 and the failure modes from REL 3 turn into architecture, and where most of the money gets spent.

Two mistakes account for most of the disappointment. The first is making compute redundant and leaving state on a single node, which produces a system that survives everything except the failure that matters. The second is building a failover path that has never run, which is a theory rather than a mechanism.

  • REL 4.1 Distribute compute across availability zones according to the flow’s target
  • REL 4.2 Make state redundant, and know where each data set is anchored
  • REL 4.3 Verify the failover path is not itself a single point of failure
  • REL 4.4 Decide deliberately whether the workload needs a second region

REL 4.1 Distribute compute across availability zones according to the flow’s target

Section titled “REL 4.1 Distribute compute across availability zones according to the flow’s target”

Risk if not established: High

Zone distribution removes an entire class of outage for roughly the cost of running the second instance. It is the highest return available in this pillar, and the return drops sharply after the first step: going from one zone to two buys far more than going from two to three.

Which flows get it comes from the ranking in REL 2.2, not from applying it everywhere by default.

The design work is in making the components distributable. Anything holding local state, anything with a fixed identity, and anything a peer discovers by address will resist. Those constraints are easier to design out at the start than to retrofit.

On STACKIT. Every region provides at least three availability zones, each with separate power, cooling and local network connectivity. In eu01 these are eu01-1, eu01-2, eu01-3, plus a metro zone eu01-m.

The Compute Engine service certificate prices these topologies directly, and the difference between one zone, a Metro Availability Zone and a two-zone system group is large enough to decide the design on its own. The figures and how to compose them are in REL 1.3, so they are not repeated here.

In Kubernetes Engine , a single node pool can span multiple zones and distributes nodes evenly across them. Two properties shape the configuration. The node maximum must be at least the number of selected zones, so a pool spanning three zones cannot have a maximum of two. And zones cannot be removed from an existing node pool: adding is supported, removing means creating a new pool and migrating the workload. Choosing the zone set is therefore a decision with a one-way component, which is unusual enough to be worth planning rather than discovering.

One caveat carried over from REL 1.3: the certificate states that several availability zones can be located in the same building. Zone redundancy addresses infrastructure failure. Physical separation across sites is what regions answer, which is REL 4.4.

Tradeoffs. Cost Optimization. Roughly doubles the compute for the first step, and adds inter-zone traffic. Performance Efficiency. Cross-zone calls cost latency, which matters for chatty components and rarely for anything else. Sustainability. Standby capacity that does no work is the sharpest conflict in SUS 4; prefer active-active where the workload allows it.

Verify. For each component on your highest-ranked flow, how many availability zones does it run in, and when was the loss of one zone last exercised rather than assumed?


REL 4.2 Make state redundant, and know where each data set is anchored

Section titled “REL 4.2 Make state redundant, and know where each data set is anchored”

Risk if not established: High

Stateless components are easy to duplicate, which is why teams do that part and stop. State is where redundancy is hard, expensive, and load-bearing.

Three decisions per data set:

Replication mode. Synchronous replication costs write latency and closes the data-loss window. Asynchronous returns the latency and reopens it. This is an RPO decision from REL 1.2, not a performance decision, and it is made per data set rather than per system. Not all data warrants a synchronous write.

Anchoring. Where can this data actually be read from? A volume that exists in one zone constrains everything that needs it, regardless of how many zones the compute spans.

Failover behaviour. Who promotes a replica, how long it takes, and what happens to in-flight writes. If the answer is “an engineer does it”, that duration belongs in your RTO.

On STACKIT. PostgreSQL Flex offers single instances and replica sets of three nodes, described as full mirrors of each other, with the three-replica type recommended for production. Transaction logs are enabled by default, which is what makes point-in-time recovery possible across snapshots. See REL 8.

Aus der STACKIT-DokuArchitektur von PostgreSQL Flex › Instanz-EbeneStand der Quelle 17.11.2025 · übernommen 05.10.2026

Auf der Instanz-Ebene definieren Sie, auf wie vielen Knoten Ihre Instanz läuft. Die Eigenschaft Typ legt fest, ob eine Instanz eine Einzelinstanz oder ein Replica-Set ist. Ein Replica-Set besteht aus 3 Knoten für die Produktionsresilienz. Alle drei Knoten sind ein vollständiger Spiegel voneinander.

Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

Replication between those three nodes is synchronous: a commit reaches the application only once the other nodes have written it, and the nodes sit in different availability zones. Both facts matter to a target. Synchronous replication across zones means a zone failure costs no committed transactions, so the data-loss side of REL 1.1 is the strong one here. It also means every write pays the latency between zones, which is the cost PERF 4 accounts for and not a setting you can trade away.

Synchronous replication settles the data-loss side of the target. The downtime side needs a second figure, what triggers a failover and how long it takes, and that one is established by rehearsing it under REL 10 rather than read off a page. A recovery-time objective that has never been checked against an observed failover is a guess.

In Kubernetes Engine, a persistent volume is anchored to one availability zone and cannot be mounted from another without migrating the storage , which is the standard behaviour for zonal block storage. The consequence is specific and easy to miss: a node pool spanning three zones does not make a stateful workload zone-redundant. The pod follows its volume. Zone redundancy for state comes from the data layer replicating, not from the scheduler.

The zone is chosen when the first pod using the claim is scheduled, not when the claim is written: the storage classes bind with WaitForFirstConsumer, so the volume is created in the zone the scheduler landed on. Placement follows scheduling once, and is fixed from then on. That is worth knowing for the first placement and no help at all afterwards.

Object Storage is the natural home for data that must outlive any single zone, and it is S3-compatible, which also serves SOV 10.

Aus der STACKIT-DokuArchitektur der Object Storage › Regionen und VerfügbarkeitszonenStand der Quelle 21.07.2026 · übernommen 05.10.2026

Die Quelle zeigt hier Bilder, die auf dieser Seite fehlen (1). Zu sehen in Architektur der Object Storage

Object Storage ist in bestimmten Regionen verfügbar. Innerhalb jeder Region werden Ihre Daten automatisch über alle drei Verfügbarkeitszonen repliziert. Diese Replikation erfolgt transparent, Sie müssen sie nicht konfigurieren.

Das folgende Diagramm zeigt die Struktur von Object Storage in einer Beispielregion (EU01):

Jede Region verfügt über eine dedizierte Endpunkt-URL. Sie müssen den Endpunkt der Region verwenden, in der sich Ihr Bucket befindet.

Alle STACKIT Object Storage-Endpunkte unterstützen die TLS 1.3-Verschlüsselung.

Buckets sind regionsspezifisch. Sie können einen Bucket nach der Erstellung nicht zwischen Regionen verschieben.

Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

Tradeoffs. Performance Efficiency. Synchronous replication is paid for on every write. Cost Optimization. Replicated storage multiplies volume and adds inter-zone traffic. Sovereignty & Compliance. Every replica is another copy of the data, inheriting its classification under SOV 3.

Verify. For each stateful component, how many zones is the data replicated across, is replication synchronous, and how long does promotion take? Which of those answers is documented rather than assumed?


REL 4.3 Verify the failover path is not itself a single point of failure

Section titled “REL 4.3 Verify the failover path is not itself a single point of failure”

Risk if not established: High

The mechanism that protects you is a component like any other. It can be misconfigured, it can hold an expired certificate, and it can depend on something that fails at the same time as the thing it was protecting.

The specific pattern to look for is a redundancy mechanism with a non-redundant control path: a single load balancer in front of a multi-zone backend, a failover script that runs on one host, a health check that queries a monitoring system in the failing zone, or a DNS change that requires access to a console you cannot reach.

Complexity is a failure mode in itself. A failover path intricate enough that nobody can explain it in a few minutes will not work at three in the morning, when the person on call did not build it. Simplicity is a reliability feature, and it is worth trading some theoretical coverage for a mechanism that people actually understand.

On STACKIT. Load balancing is the usual entry point to a redundant backend, and its own placement and redundancy belong in the analysis rather than being assumed. Check its service certificate alongside the ones for what sits behind it, exactly as in REL 1.3.

In Kubernetes Engine, the Metro Availability Zone provides automatic failover for applications that are not themselves resilient, while a multi-zone node pool delivers availability for workloads designed for it. Those are two different mechanisms with different assumptions, and choosing between them is worth doing explicitly rather than by default.

Aus der STACKIT-DokuTopologies › Single availability zone and metro availability zoneStand der Quelle 19.08.2026 · übernommen 05.10.2026

The Metro AZ is a special STACKIT offering that is actually a High Availability Zone spanning the other three AZs (eu01-1 to eu01-03, called Single AZs). It is designed for applications that cannot achieve resilience on their own - in case of a failure of one of the Single AZs an application placed in the Metro AZ is started in one of the other Single AZs. You can learn more about this concept here: Block Storage service plans.

As described above, the built-in Kubernetes mechanisms and its ability to span multiple AZs make it very failure-resistant already. You do not have to use the Metro AZ in SKE to achieve High Availability if the node pools are configured correctly to run in multiple AZs.

Was ist das?

Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.

Access to the control plane during an incident is part of this. If recovery requires the STACKIT Portal, the API or the CLI, then authentication is on your recovery path, which is the recovery-dependency point from REL 3.3.

Tradeoffs. Operational Excellence. Every failover mechanism needs testing and tuning, and a neglected one is worse than none because it creates false confidence.

Verify. Draw the path a request takes during a failover. Which components on that path are not themselves redundant, and which of them are needed to trigger the failover at all?


REL 4.4 Decide deliberately whether the workload needs a second region

Section titled “REL 4.4 Decide deliberately whether the workload needs a second region”

Risk if not established: Medium

Multi-region is the most expensive step in this pillar and the one most often taken for the wrong reason. It protects against a whole-region loss, which is rare, and it introduces cross-region data consistency, which is permanent.

Make it an explicit decision with three inputs: whether your RTO and RPO can be met within one region, whether the data can legally and practically live in the second, and whether anyone will maintain the second environment well enough for it to work when needed. A cold standby that has drifted for a year is not a recovery capability.

Before reaching for it, check the cheaper options: a second zone, backups held outside the primary region under REL 8.2, or an accepted longer RTO. Most workloads that ask for multi-region actually want one of those.

On STACKIT. There are two regions, eu01 in Germany and eu02 in Austria, described in regions and availability zones .

Both sit inside EU jurisdiction, which makes cross-region recovery materially simpler here than on platforms whose second region may not. It removes a constraint rather than a decision: SOV 2 still requires the placement to be recorded against your data classification, and the sovereignty tier from SOV 1 still governs.

The platform documents the building blocks for a second region rather than a managed cross-region failover pattern, so the orchestration, the data replication and the promotion procedure are yours to design and operate. That is the usual division for infrastructure services, and it is worth planning for explicitly rather than expecting a switch.

Object Storage is not the exception to that. It does not replicate across regions, and no such feature is announced, so a second region means you copy the objects yourself, on a schedule you choose, and the recovery point for a region loss is exactly that schedule’s interval. Within a region it is spread across availability zones by construction rather than by configuration, which is a different guarantee and covers a different failure.

Tradeoffs. Cost Optimization. Can approach doubling the workload cost for a failure mode that may never occur. Operational Excellence. A second environment to patch, monitor and keep from drifting. Sustainability. Idle standby capacity, SUS 4.

Verify. Can your RTO and RPO be met inside a single region? If yes, what is the second region for, and who agreed that the reason justifies its cost?


  • REL 1 Reliability targets, particularly the composition in REL 1.3
  • REL 3 Failure modes, which redundancy is one response to
  • REL 8 Backup and restore, which covers what redundancy cannot
  • REL 10 Health model and testing, which is how you find out whether any of this works
  • SOV 2 Placement and residency, which constrains where redundancy may be placed
  • Reliability tradeoffs