Reliability: tradeoffs
Last updated on
Reliability is the pillar most often pursued without acknowledging its price, because the price arrives later and lands on someone else’s budget. This page states what it costs, so that the decision to pay is a decision.
Against Cost Optimization
Section titled “Against Cost Optimization”Redundancy against spend. Redundancy means running capacity that produces nothing until something fails. Multi-zone replication multiplies storage and adds inter-zone traffic. A warm standby in a second region can approach the cost of the primary while serving no requests at all.
The cost does not scale linearly with the benefit. Moving from a single instance to two across zones removes an entire class of outage for roughly double the compute. Moving from two to three buys considerably less. Moving to a second region can double the total workload cost for a failure mode that may never occur in the system’s lifetime.
How to resolve it: this is exactly what REL 1 and REL 2 exist for. Priced against a
business-agreed RTO and a ranked list of flows, the argument becomes arithmetic instead of
opinion. Without those inputs, the reliability-versus-cost discussion is two people asserting
preferences.
Against Operational Excellence
Section titled “Against Operational Excellence”Every reliability mechanism is something to operate. Failover logic must be tested; standby environments drift from primaries; circuit breakers need thresholds that someone tunes; backup jobs fail silently. The mechanisms that protect you also demand attention, and neglected ones are worse than absent ones because they create false confidence.
There is a real inversion point. A sufficiently intricate resilience design reduces reliability, because nobody understands it well enough to operate it under pressure. This is principle 3 as a concrete cost.
How to resolve it: prefer mechanisms the platform operates over ones you operate. Managed service failover you configure once beats orchestration you maintain. Where you must build it, budget for the exercising, not just the building.
Against Performance Efficiency
Section titled “Against Performance Efficiency”Usually aligned, capacity headroom serves both, but they diverge in specific places:
- Synchronous cross-zone replication costs write latency to buy durability. Asynchronous replication returns the latency and reintroduces the data-loss window.
- Retries improve success rates and increase load, precisely when the system is already struggling. A retry policy without backoff and a circuit breaker converts a slow dependency into an outage.
- Timeouts protect the caller by abandoning work the callee may have completed.
- Health checks and quorum protocols consume real capacity in large deployments.
How to resolve it: treat replication mode as an RPO decision, not a performance decision, and make it per data set rather than per system. Not all data warrants a synchronous write.
Against Sustainability
Section titled “Against Sustainability”Idle standby capacity consumes resources for no delivered work, the sharpest divergence between Cost Optimization and Sustainability in the framework, since a reservation can be financially efficient and physically wasteful at the same time.
How to resolve it: prefer active-active over active-passive where the workload permits, so that redundant capacity does useful work. Where standby is unavoidable, keep it minimal and scale it on failover rather than running a full mirror.
Against Security
Section titled “Against Security”Mostly complementary: segmentation limits both blast radius and attack surface. Two frictions:
- Backups are copies of your data, inheriting its classification and multiplying the places it
must be protected.
REL 8andSEC 7have to be read together. - Disaster recovery procedures need elevated access, and emergency access paths are attractive targets. A break-glass procedure that bypasses normal controls is a reliability mechanism and a security liability at once.
Against Sovereignty & Compliance
Section titled “Against Sovereignty & Compliance”Cross-region redundancy is a data placement decision. Replicating to a second region moves
data across a border. Within STACKIT, eu01 in Germany, eu02 in Austria, both remain in EU
jurisdiction, which makes this materially simpler than on providers whose second region may not.
It is still a decision that SOV 2 requires you to make explicitly rather than inherit from a
replication default.
Backups carry the same constraint, including their retention: a backup in a location or under a retention rule your classification does not permit is a compliance finding regardless of how well it serves recovery.
Reading the tradeoffs
Section titled “Reading the tradeoffs”Every question page carries the tradeoffs of its individual best practices. This page is the pillar-level view.
- Design principles
- Overview: the questions this pillar asks
- Cost Optimization tradeoffs: the same conflicts from the other side