---
id: REL01
pillar: reliability
title: REL 1. How do you derive reliability targets from business impact?
description: Availability, RTO and RPO come from what an outage costs the business, not from what the platform happens to offer. How to derive them and check them.
status: draft
services: [compute-engine, kubernetes-engine, postgresql-flex]
sidebar:
  order: 10
  label: Reliability targets
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/rel-01-reliability-targets/"
source_file: "docs/architecture/pillars/reliability/rel-01-reliability-targets.mdx"
---

Most workloads have no stated reliability target. They have an implicit one, assembled from
whatever the team happened to build, and nobody discovers what it is until the first serious
outage. At that point two things become clear at once: the business expected more than the
architecture delivers, and nobody can say what "more" would have cost.

A stated target ends both arguments before they start. It tells you which redundancy is worth
paying for, it tells a cost review which savings are not available, and it turns "is this reliable
enough" from an opinion into a comparison.

## Best practices

- [`REL 1.1`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-11-quantify-what-an-outage-and-what-data-loss-cost-the-business) Quantify what an outage and what data loss cost the business
- [`REL 1.2`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-12-set-availability-rto-and-rpo-per-critical-flow-rather-than-per-workload) Set availability, RTO and RPO per critical flow rather than per workload
- [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on) Check every target against the published availability of the services it depends on
- [`REL 1.4`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-14-have-each-target-agreed-and-recorded-by-someone-accountable-for-the-outcome) Have each target agreed and recorded by someone accountable for the outcome
- [`REL 1.5`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-15-keep-your-internal-objective-stricter-than-any-commitment-you-make-to-others) Keep your internal objective stricter than any commitment you make to others

---

## REL 1.1 Quantify what an outage and what data loss cost the business

**Risk if not established:** Medium

Reliability is bought. Until someone knows the price of not having it, every architectural
decision about redundancy is a guess wearing the clothes of engineering.

The figure you need is a cost per unit of downtime, and it is rarely just lost revenue.
Contractual penalties, staff who cannot work, recovery labour, regulatory notification duties, and
the cost of re-acquiring a customer who left all belong in it. For some workloads the honest
answer is that an hour of downtime costs almost nothing, and that is a useful answer: it tells you
to stop over-engineering.

Data loss needs its own number, because it behaves differently. Downtime cost accumulates with
time; data loss cost is often a step function. Losing five minutes of transactions may be an
inconvenience, while losing an hour may be unrecoverable. The shape of that curve determines
whether synchronous replication is worth its latency.

The common failure is to skip this because the business conversation is harder than the technical
one. What follows is a workload over-engineered where the team found the problem interesting and
under-engineered everywhere else, with no way to tell which is which.

**On STACKIT.** Nothing on the platform produces this number for you. It comes from the business,
and no cloud provider can supply it.

Two things the platform does help with. The <LinkChip href="https://docs.stackit.cloud/platform/cost-and-billing/cost-dashboard/">Cost Dashboard</LinkChip>
gives you the other side of the equation, which is what the reliability you are considering would
actually cost to run. And STACKIT publishes a
<LinkChip href="https://stackit.com/en/gtc/service-certificates">service certificate</LinkChip> for every service, which
tells you what a given level of reliability costs in architecture rather than in guesswork. See
[`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on).

**Tradeoffs.** None between pillars. The cost is stakeholder time, mostly outside engineering, and
the answer is frequently uncomfortable. There is no way to arrive at a defensible target without it.

**Verify.** For your highest-ranked flow, what does one hour of unavailability cost, what does one
hour of lost data cost, and who produced those figures?

---

## REL 1.2 Set availability, RTO and RPO per critical flow rather than per workload

**Risk if not established:** Medium

A workload-level target is wrong for at least one part of the workload. The checkout path and the
monthly report have legitimately different requirements, and a single number applied to both
either over-protects the report or under-protects the checkout.

Three separate figures are needed, and they are frequently confused:

- **Availability** is the fraction of time the flow works. It sets your redundancy.
- **RTO**, the recovery time objective, is how long the flow may stay broken. It sets your
  recovery procedure and how often you rehearse it.
- **RPO**, the recovery point objective, is how much recent data may be lost. It sets your backup
  and replication frequency.

All three are business inputs, decided before the architecture rather than derived from it. A
target that was calculated from what the current design achieves is not a target; it is a
description.

The ranking of flows comes from [`REL 2`](/architecture/pillars/reliability/rel-02-critical-flows/), and doing it once serves several pillars: reliability
investment, cost scrutiny under [`COST 7`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/), and performance targets under [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) all follow the
same order.

**On STACKIT.** Per-flow targets are the natural unit here, because STACKIT's own commitments are
made per service rather than per workload. A flow that touches Kubernetes Engine, PostgreSQL Flex
and Object Storage inherits three separate figures, and only the flow-level view composes them.
See [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on).

**Tradeoffs.** More targets to agree, track and review. **Cost Optimization** benefits: per-flow
targets are what allow you to spend less on the flows that need less, which a single
workload-level number never permits.

**Verify.** For each critical flow, what are the availability target, the RTO and the RPO, and
which flows differ from each other?

---

## REL 1.3 Check every target against the published availability of the services it depends on

**Risk if not established:** Medium

A target above what your dependencies support is a wish. Since dependencies on a flow are usually
serial, their availabilities multiply: four services each committed to 99.9% give you roughly
99.6% before your own code has failed once.

So the check runs in two directions. Compose the published figures for everything the flow touches
and compare the result against your target. Where the composition falls short, you have three
honest options: add redundancy to raise the ceiling, remove a dependency from the critical path,
or lower the target. Deciding to do nothing is a fourth option, and it needs to be recorded as an
accepted risk rather than left implicit.

Read what the numbers actually mean, not just their size. Availability definitions differ in ways
that matter, and the exclusions are usually where the surprises live.

**On STACKIT.** STACKIT publishes a
<LinkChip href="https://stackit.com/en/gtc/service-certificates">service certificate</LinkChip> for every service, each
stating an agreed availability. Every figure this calculation needs is therefore on one page and
stated per service, which is what makes it a calculation rather than an estimate.

The Compute Engine certificate also shows how much the deployment shape changes the figure, which
is the lever you actually control:

| Deployment | Agreed availability, calendar month average |
|---|---|
| Single VM in one availability zone | 99.5% |
| VM in a Metro Availability Zone | 99.8% |
| System group, two VMs in two availability zones in one region | 99.9% for at least one VM |

Reading any provider's availability figure correctly means checking three things, and the STACKIT
certificates state all three plainly rather than leaving them to be inferred.

**What counts as unavailable.** For Compute Engine it is the loss of external connectivity. Every
provider defines availability narrowly like this, because it has to be measurable without knowing
what your application does. It follows that your health model under [`REL 10`](/architecture/pillars/reliability/rel-10-health-model-and-testing/) should be stricter
than the SLA definition: you want to detect the degradation that a connectivity check was never
designed to see.

**Where one service ends and the next begins.** Block Storage carries its own certificate, so a
storage fault is measured against that document rather than against the VM's. Compose both
certificates when a flow depends on both. This is the normal shape of a service catalogue and the
reason per-service certificates are useful in the first place.

**What a zone boundary buys.** Each availability zone has separate power, cooling and local
network connectivity. The certificate also states that several zones can be located in the same
building. Zone redundancy therefore addresses infrastructure failure, while separation across sites
is a different question, and regions are what answer it. The design consequence belongs in
[`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/).

Finally, check which side of the shared responsibility line each part of your RPO sits on. For
Compute Engine, backup and recovery are the customer's responsibility, which is the usual division
for infrastructure services. STACKIT offers <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/">Server Backup
Management</LinkChip> as a
managed option; the point is that your RPO should name which of the two it relies on rather than
assume. See [`REL 8`](/architecture/pillars/reliability/rel-08-backup-and-restore/).

Read the published figures per region rather than assuming one covers both `eu01` and `eu02`, and
have a target that leans on the distinction name the region it was read for. Whether Metro
Availability Zones exist in both bears on the same decision and is covered under [`REL 4.4`](/architecture/pillars/reliability/rel-04-redundancy/#rel-44-decide-deliberately-whether-the-workload-needs-a-second-region).

**Tradeoffs.** **Cost Optimization.** Raising the achievable ceiling means redundancy, and the
cost does not scale linearly with the benefit: moving from a single VM to a two-zone system group
buys an entire class of outage for roughly double the compute, while the next increment buys much
less.

**Verify.** For your most critical flow, which STACKIT services does it depend on, what does each
certificate commit to, and what does the composition of those figures allow compared with your
stated target?

---

## REL 1.4 Have each target agreed and recorded by someone accountable for the outcome

**Risk if not established:** Medium

An unowned target does not survive contact with a budget. Under cost pressure, backup retention
gets shortened, a standby gets downsized, a third replica is dropped. Each change is individually
defensible, none is announced as a reduction in reliability, and the workload drifts away from
what the business assumed without anyone deciding that it should.

A recorded target with a name against it converts each of those into a decision somebody has to
make explicitly. That is its main purpose. Record what was agreed, who agreed it, when, and what
the agreement was based on, because the reasoning ages faster than the number.

Set a trigger for revisiting rather than a calendar reminder nobody honours: a material change in
the business, a new regulatory obligation, or an architecture change that alters the dependency
composition from [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on).

**On STACKIT.** No platform feature applies. This is an organizational practice, and the record
belongs wherever your architecture decisions live rather than in the cloud environment.

**Tradeoffs.** **Operational Excellence.** Another artefact to keep current, and a stale target is
worse than none because it carries false authority.

**Verify.** Who agreed the RTO for your most critical flow, on what date, and where is that
recorded?

---

## REL 1.5 Keep your internal objective stricter than any commitment you make to others

**Risk if not established:** Medium

If your internal objective equals your external commitment, you have no margin. You discover you
are in breach at the same moment your customer does, and every incident becomes a contractual
event.

Set the internal objective tighter, and treat the gap as your working room. Crossing it is a
signal to act; crossing the external one is a failure. The size of the gap is a judgement about
how quickly you detect and recover, which means it should be informed by the RTO from [`REL 1.2`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-12-set-availability-rto-and-rpo-per-critical-flow-rather-than-per-workload)
rather than picked round.

Keep the two labelled distinctly. An objective you set for yourself and a commitment someone can
hold you to are different in kind, and conflating them either makes you over-cautious about
internal goals or careless about external ones.

**On STACKIT.** The service certificates are STACKIT's commitments to you, not objectives. Your
commitments to your own customers sit on top of them and must absorb your own failure modes as
well, which is why they can never simply restate the platform figure.

**Tradeoffs.** Little. A stricter internal objective may trigger work that a purely contractual
reading would not require, which is the point of having one.

**Verify.** What is the difference between your internal availability objective and any commitment
you have made externally, and what does that difference buy you in response time?

---

## Related

- [`REL 2`](/architecture/pillars/reliability/rel-02-critical-flows/) Critical flows, which supplies the ranking these targets attach to
- [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/) Redundancy, where the composition from [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on) turns into architecture
- [`REL 8`](/architecture/pillars/reliability/rel-08-backup-and-restore/) Backup and restore, which the RPO drives
- [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/) Disaster recovery, which the RTO drives
- [Reliability tradeoffs](/architecture/pillars/reliability/tradeoffs/), particularly the conflict with Cost Optimization
- <LinkChip href="https://stackit.com/en/gtc/service-certificates">STACKIT service certificates</LinkChip>
