---
id: REL04
pillar: reliability
title: REL 4. How do you eliminate single points of failure?
description: Redundancy has to cover state as well as compute, and the failover path must not itself be the weak point. What STACKIT zones and regions actually give you.
status: draft
services: [compute-engine, kubernetes-engine, postgresql-flex, object-storage]
sidebar:
  order: 13
  label: Redundancy
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/rel-04-redundancy/"
source_file: "docs/architecture/pillars/reliability/rel-04-redundancy.mdx"
---

This is where the targets from [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) and the failure modes from [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) turn into architecture,
and where most of the money gets spent.

Two mistakes account for most of the disappointment. The first is making compute redundant and
leaving state on a single node, which produces a system that survives everything except the
failure that matters. The second is building a failover path that has never run, which is a theory
rather than a mechanism.

## Best practices

- [`REL 4.1`](/architecture/pillars/reliability/rel-04-redundancy/#rel-41-distribute-compute-across-availability-zones-according-to-the-flows-target) Distribute compute across availability zones according to the flow's target
- [`REL 4.2`](/architecture/pillars/reliability/rel-04-redundancy/#rel-42-make-state-redundant-and-know-where-each-data-set-is-anchored) Make state redundant, and know where each data set is anchored
- [`REL 4.3`](/architecture/pillars/reliability/rel-04-redundancy/#rel-43-verify-the-failover-path-is-not-itself-a-single-point-of-failure) Verify the failover path is not itself a single point of failure
- [`REL 4.4`](/architecture/pillars/reliability/rel-04-redundancy/#rel-44-decide-deliberately-whether-the-workload-needs-a-second-region) Decide deliberately whether the workload needs a second region

---

## REL 4.1 Distribute compute across availability zones according to the flow's target

**Risk if not established:** High

Zone distribution removes an entire class of outage for roughly the cost of running the second
instance. It is the highest return available in this pillar, and the return drops sharply after
the first step: going from one zone to two buys far more than going from two to three.

Which flows get it comes from the ranking in [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume), not from applying it everywhere by
default.

The design work is in making the components distributable. Anything holding local state, anything
with a fixed identity, and anything a peer discovers by address will resist. Those constraints are
easier to design out at the start than to retrofit.

**On STACKIT.** Every <LinkChip href="https://docs.stackit.cloud/platform/regions/">region</LinkChip> provides at least
three availability zones, each with separate power, cooling and local network connectivity. In
`eu01` these are `eu01-1`, `eu01-2`, `eu01-3`, plus a metro zone `eu01-m`.

The Compute Engine <LinkChip href="https://stackit.com/en/gtc/service-certificates">service certificate</LinkChip> prices
these topologies directly, and the difference between one zone, a Metro Availability Zone and a
two-zone system group is large enough to decide the design on its own. The figures and how to
compose them are in [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on), so they are not repeated here.

In <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/operations/topologies/">Kubernetes
Engine</LinkChip>, a
single node pool can span multiple zones and distributes nodes evenly across them. Two properties
shape the configuration. The node maximum must be at least the number of
selected zones, so a pool spanning three zones cannot have a maximum of two. And **zones cannot be
removed from an existing node pool**: adding is supported, removing means creating a new pool and
migrating the workload. Choosing the zone set is therefore a decision with a one-way component,
which is unusual enough to be worth planning rather than discovering.

One caveat carried over from [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on): the certificate states that several availability zones can
be located in the same building. Zone redundancy addresses infrastructure failure. Physical
separation across sites is what regions answer, which is [`REL 4.4`](/architecture/pillars/reliability/rel-04-redundancy/#rel-44-decide-deliberately-whether-the-workload-needs-a-second-region).

**Tradeoffs.** **Cost Optimization.** Roughly doubles the compute for the first step, and adds
inter-zone traffic. **Performance Efficiency.** Cross-zone calls cost latency, which matters for
chatty components and rarely for anything else. **Sustainability.** Standby capacity that does no
work is the sharpest conflict in [`SUS 4`](/architecture/pillars/sustainability/sus-04-utilization-density/); prefer active-active where the workload allows it.

**Verify.** For each component on your highest-ranked flow, how many availability zones does it
run in, and when was the loss of one zone last exercised rather than assumed?

---

## REL 4.2 Make state redundant, and know where each data set is anchored

**Risk if not established:** High

Stateless components are easy to duplicate, which is why teams do that part and stop. State is
where redundancy is hard, expensive, and load-bearing.

Three decisions per data set:

**Replication mode.** Synchronous replication costs write latency and closes the data-loss window.
Asynchronous returns the latency and reopens it. This is an RPO decision from [`REL 1.2`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-12-set-availability-rto-and-rpo-per-critical-flow-rather-than-per-workload), not a
performance decision, and it is made per data set rather than per system. Not all data warrants a
synchronous write.

**Anchoring.** Where can this data actually be read from? A volume that exists in one zone
constrains everything that needs it, regardless of how many zones the compute spans.

**Failover behaviour.** Who promotes a replica, how long it takes, and what happens to in-flight
writes. If the answer is "an engineer does it", that duration belongs in your RTO.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/basics/architecture-of-postgresql-flex/">PostgreSQL
Flex</LinkChip>
offers single instances and replica sets of three nodes, described as full mirrors of each other,
with the three-replica type recommended for production. Transaction logs are
enabled by default, which is what makes point-in-time recovery possible across snapshots. See `REL
8`.

> From the STACKIT docs: [Architecture › Instance level](https://docs.stackit.cloud/products/databases/postgresql-flex/basics/architecture-of-postgresql-flex/#instance-level) (Source updated 19.03.2026, copied 05.10.2026)

On the instance level, you define how many nodes your instance runs on. The **type** property defines whether an instance is a **single instance** or a **replica set**. A **replica set** consists of 3 nodes for production resilience. All three nodes are a full mirror of each other.

Replication between those three nodes is **synchronous**: a commit reaches the application only
once the other nodes have written it, and the nodes sit in different availability zones. Both
facts matter to a target. Synchronous replication across zones means a zone failure costs no
committed transactions, so the data-loss side of [`REL 1.1`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-11-quantify-what-an-outage-and-what-data-loss-cost-the-business) is the strong one here. It also means
every write pays the latency between zones, which is the cost [`PERF 4`](/architecture/pillars/performance-efficiency/perf-04-data-design/) accounts for and not a
setting you can trade away.

Synchronous replication settles the data-loss side of the target. The downtime side needs a
second figure, what triggers a failover and how long it takes, and that one is established by
rehearsing it under [`REL 10`](/architecture/pillars/reliability/rel-10-health-model-and-testing/) rather than read off a page. A recovery-time objective that has never
been checked against an observed failover is a guess.

In Kubernetes Engine, a persistent volume is anchored to one availability zone and cannot be
mounted from another without <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/migrate-storage-to-another-az/">migrating the
storage</LinkChip>,
which is the standard behaviour for zonal block storage. The consequence is specific and easy to
miss: a node pool spanning three zones does not make a stateful workload zone-redundant. The pod
follows its volume. Zone redundancy for state comes from the data layer replicating, not from the
scheduler.

The zone is chosen when the first pod using the claim is scheduled, not when the claim is written:
the storage classes bind with `WaitForFirstConsumer`, so the volume is created in the zone the
scheduler landed on. Placement follows scheduling once, and is fixed from then on. That is worth
knowing for the first placement and no help at all afterwards.

<LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/">Object Storage</LinkChip> is the natural home
for data that must outlive any single zone, and it is S3-compatible, which also serves [`SOV 10`](/architecture/pillars/sovereignty/sov-10-open-interfaces/).

> From the STACKIT docs: [Architecture of Object Storage › Regions and availability zones](https://docs.stackit.cloud/products/storage/object-storage/basics/architecture/#regions-and-availability-zones) (Source updated 21.07.2026, copied 05.10.2026)

Object Storage is available in dedicated regions. Within each region, your data is automatically replicated across all three **Availability Zones**. This replication happens transparently, you do not need to configure it.

The following diagram shows the structure of Object Storage in an example region (EU01):

Each region has a dedicated **endpoint URL**. You need to use the endpoint of the region where your bucket resides.

| Region | Endpoint URL |
| --- | --- |
| EU01 | `https://object.storage.eu01.onstackit.cloud` |
| EU02 | `https://object.storage.eu02.onstackit.cloud` |

All STACKIT Object Storage endpoints support TLS 1.3 encryption.

Buckets are region-specific. You can not move a bucket between regions after creation.

**Tradeoffs.** **Performance Efficiency.** Synchronous replication is paid for on every write.
**Cost Optimization.** Replicated storage multiplies volume and adds inter-zone traffic.
**Sovereignty & Compliance.** Every replica is another copy of the data, inheriting its
classification under [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/).

**Verify.** For each stateful component, how many zones is the data replicated across, is
replication synchronous, and how long does promotion take? Which of those answers is documented
rather than assumed?

---

## REL 4.3 Verify the failover path is not itself a single point of failure

**Risk if not established:** High

The mechanism that protects you is a component like any other. It can be misconfigured, it can
hold an expired certificate, and it can depend on something that fails at the same time as the
thing it was protecting.

The specific pattern to look for is a redundancy mechanism with a non-redundant control path: a
single load balancer in front of a multi-zone backend, a failover script that runs on one host, a
health check that queries a monitoring system in the failing zone, or a DNS change that requires
access to a console you cannot reach.

Complexity is a failure mode in itself. A failover path intricate enough that nobody can explain
it in a few minutes will not work at three in the morning, when the person on call did not build
it. Simplicity is a reliability feature, and it is worth trading some theoretical coverage for a
mechanism that people actually understand.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/">Load
balancing</LinkChip> is
the usual entry point to a redundant backend, and its own placement and redundancy belong in the
analysis rather than being assumed. Check its service certificate alongside the ones for what sits
behind it, exactly as in [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on).

In Kubernetes Engine, the Metro Availability Zone provides automatic failover for applications
that are not themselves resilient, while a multi-zone node pool delivers
availability for workloads designed for it. Those are two different mechanisms with different
assumptions, and choosing between them is worth doing explicitly rather than by default.

> From the STACKIT docs: [Topologies › Single availability zone and metro availability zone](https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/operations/topologies/#single-availability-zone-and-metro-availability-zone) (Source updated 19.08.2026, copied 05.10.2026)

The Metro AZ is a special STACKIT offering that is actually a High Availability Zone spanning the other three AZs (eu01-1 to eu01-03, called Single AZs). It is designed for applications that cannot achieve resilience on their own - in case of a failure of one of the Single AZs an application placed in the Metro AZ is started in one of the other Single AZs. You can learn more about this concept here: [Block Storage service plans](https://docs.stackit.cloud/products/storage/block-storage/basics/service-plans/).

As described above, the built-in Kubernetes mechanisms and its ability to span multiple AZs make it very failure-resistant already. You do not have to use the Metro AZ in SKE to achieve High Availability if the node pools are configured correctly to run in multiple AZs.

Access to the control plane during an incident is part of this. If recovery requires the STACKIT
Portal, the API or the CLI, then authentication is on your recovery path, which is the
recovery-dependency point from [`REL 3.3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-33-include-the-failure-modes-of-your-dependencies-not-only-of-your-own-code).

**Tradeoffs.** **Operational Excellence.** Every failover mechanism needs testing and tuning, and
a neglected one is worse than none because it creates false confidence.

**Verify.** Draw the path a request takes during a failover. Which components on that path are not
themselves redundant, and which of them are needed to trigger the failover at all?

---

## REL 4.4 Decide deliberately whether the workload needs a second region

**Risk if not established:** Medium

Multi-region is the most expensive step in this pillar and the one most often taken for the wrong
reason. It protects against a whole-region loss, which is rare, and it introduces cross-region
data consistency, which is permanent.

Make it an explicit decision with three inputs: whether your RTO and RPO can be met within one
region, whether the data can legally and practically live in the second, and whether anyone will
maintain the second environment well enough for it to work when needed. A cold standby that has
drifted for a year is not a recovery capability.

Before reaching for it, check the cheaper options: a second zone, backups held outside the primary
region under [`REL 8.2`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-82-hold-copies-where-a-single-failure-cannot-destroy-both), or an accepted longer RTO. Most workloads that ask for multi-region
actually want one of those.

**On STACKIT.** There are two regions, `eu01` in Germany and `eu02` in Austria, described in
<LinkChip href="https://docs.stackit.cloud/platform/regions/">regions and availability zones</LinkChip>.

Both sit inside EU jurisdiction, which makes cross-region recovery materially simpler here than on
platforms whose second region may not. It removes a constraint rather than a decision:
[`SOV 2`](/architecture/pillars/sovereignty/sov-02-placement-and-residency/) still requires the placement to be recorded against your data classification, and the
sovereignty tier from [`SOV 1`](/architecture/pillars/sovereignty/sov-01-sovereignty-tier/) still governs.

The platform documents the building blocks for a second region rather than a managed
cross-region failover pattern, so the orchestration, the data replication and the promotion
procedure are yours to design and operate. That is the usual division for infrastructure services,
and it is worth planning for explicitly rather than expecting a switch.

Object Storage is not the exception to that. It does not replicate across regions, and no such
feature is announced, so a second region means you copy the objects yourself, on a schedule you
choose, and the recovery point for a region loss is exactly that schedule's interval. Within a
region it is spread across availability zones by construction rather than by configuration, which
is a different guarantee and covers a different failure.

**Tradeoffs.** **Cost Optimization.** Can approach doubling the workload cost for a failure mode
that may never occur. **Operational Excellence.** A second environment to patch, monitor and keep
from drifting. **Sustainability.** Idle standby capacity, [`SUS 4`](/architecture/pillars/sustainability/sus-04-utilization-density/).

**Verify.** Can your RTO and RPO be met inside a single region? If yes, what is the second region
for, and who agreed that the reason justifies its cost?

---

## Related

- [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) Reliability targets, particularly the composition in [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on)
- [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) Failure modes, which redundancy is one response to
- [`REL 8`](/architecture/pillars/reliability/rel-08-backup-and-restore/) Backup and restore, which covers what redundancy cannot
- [`REL 10`](/architecture/pillars/reliability/rel-10-health-model-and-testing/) Health model and testing, which is how you find out whether any of this works
- [`SOV 2`](/architecture/pillars/sovereignty/sov-02-placement-and-residency/) Placement and residency, which constrains where redundancy may be placed
- <LinkChip href="/architecture/pillars/reliability/tradeoffs/">Reliability tradeoffs</LinkChip>
