---
type: tradeoffs
pillar: reliability
code: REL
title: "Reliability: tradeoffs"
description: What pursuing reliability costs cost optimization, operations, performance, sustainability, security and sovereignty.
status: draft
sidebar:
  label: "Tradeoffs"
  order: 2
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/tradeoffs/"
source_file: "docs/architecture/pillars/reliability/tradeoffs.mdx"
---

Reliability is the pillar most often pursued without acknowledging its price, because the price
arrives later and lands on someone else's budget. This page states what it costs, so that the
decision to pay is a decision.

---

## Against Cost Optimization

**Redundancy against spend.** Redundancy means running capacity that produces
nothing until something fails. Multi-zone replication multiplies storage and adds inter-zone
traffic. A warm standby in a second region can approach the cost of the primary while serving no
requests at all.

The cost does not scale linearly with the benefit. Moving from a single instance to two across
zones removes an entire class of outage for roughly double the compute. Moving from two to three
buys considerably less. Moving to a second region can double the total workload cost for a failure
mode that may never occur in the system's lifetime.

**How to resolve it:** this is exactly what [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) and [`REL 2`](/architecture/pillars/reliability/rel-02-critical-flows/) exist for. Priced against a
business-agreed RTO and a ranked list of flows, the argument becomes arithmetic instead of
opinion. Without those inputs, the reliability-versus-cost discussion is two people asserting
preferences.

## Against Operational Excellence

Every reliability mechanism is something to operate. Failover logic must be tested; standby
environments drift from primaries; circuit breakers need thresholds that someone tunes; backup
jobs fail silently. The mechanisms that protect you also demand attention, and neglected ones are
worse than absent ones because they create false confidence.

There is a real inversion point. A sufficiently intricate resilience design reduces reliability,
because nobody understands it well enough to operate it under pressure. This is [principle
3](/architecture/pillars/reliability/principles/) as a concrete cost.

**How to resolve it:** prefer mechanisms the platform operates over ones you operate. Managed
service failover you configure once beats orchestration you maintain. Where you must build it,
budget for the exercising, not just the building.

## Against Performance Efficiency

Usually aligned, capacity headroom serves both, but they diverge in specific places:

- **Synchronous cross-zone replication** costs write latency to buy durability. Asynchronous
  replication returns the latency and reintroduces the data-loss window.
- **Retries** improve success rates and increase load, precisely when the system is already
  struggling. A retry policy without backoff and a circuit breaker converts a slow dependency into
  an outage.
- **Timeouts** protect the caller by abandoning work the callee may have completed.
- **Health checks and quorum protocols** consume real capacity in large deployments.

**How to resolve it:** treat replication mode as an RPO decision, not a performance decision, and
make it per data set rather than per system. Not all data warrants a synchronous write.

## Against Sustainability

Idle standby capacity consumes resources for no delivered work, the sharpest divergence between
Cost Optimization and Sustainability in the framework, since a reservation can be financially
efficient and physically wasteful at the same time.

**How to resolve it:** prefer active-active over active-passive where the workload permits, so
that redundant capacity does useful work. Where standby is unavoidable, keep it minimal and scale
it on failover rather than running a full mirror.

## Against Security

Mostly complementary: segmentation limits both blast radius and attack surface. Two frictions:

- **Backups are copies of your data**, inheriting its classification and multiplying the places it
  must be protected. [`REL 8`](/architecture/pillars/reliability/rel-08-backup-and-restore/) and [`SEC 7`](/architecture/pillars/security/sec-07-encryption/) have to be read together.
- **Disaster recovery procedures need elevated access**, and emergency access paths are attractive
  targets. A break-glass procedure that bypasses normal controls is a reliability mechanism and a
  security liability at once.

## Against Sovereignty & Compliance

**Cross-region redundancy is a data placement decision.** Replicating to a second region moves
data across a border. Within STACKIT, `eu01` in Germany, `eu02` in Austria, both remain in EU
jurisdiction, which makes this materially simpler than on providers whose second region may not.
It is still a decision that [`SOV 2`](/architecture/pillars/sovereignty/sov-02-placement-and-residency/) requires you to make explicitly rather than inherit from a
replication default.

Backups carry the same constraint, including their retention: a backup in a location or under a
retention rule your classification does not permit is a compliance finding regardless of how well
it serves recovery.

---

## Reading the tradeoffs

Every question page carries the tradeoffs of its individual best practices. This page is the
pillar-level view.

- <LinkChip href="/architecture/pillars/reliability/principles/">Design principles</LinkChip>
- [Overview](/architecture/pillars/reliability/): the questions this pillar asks
- [Cost Optimization tradeoffs](/architecture/pillars/cost-optimization/tradeoffs/): the same conflicts from the
  other side
