REL 1. How do you derive reliability targets from business impact?
Last updated on
Most workloads have no stated reliability target. They have an implicit one, assembled from whatever the team happened to build, and nobody discovers what it is until the first serious outage. At that point two things become clear at once: the business expected more than the architecture delivers, and nobody can say what “more” would have cost.
A stated target ends both arguments before they start. It tells you which redundancy is worth paying for, it tells a cost review which savings are not available, and it turns “is this reliable enough” from an opinion into a comparison.
Best practices
Section titled “Best practices”REL 1.1Quantify what an outage and what data loss cost the businessREL 1.2Set availability, RTO and RPO per critical flow rather than per workloadREL 1.3Check every target against the published availability of the services it depends onREL 1.4Have each target agreed and recorded by someone accountable for the outcomeREL 1.5Keep your internal objective stricter than any commitment you make to others
REL 1.1 Quantify what an outage and what data loss cost the business
Section titled “REL 1.1 Quantify what an outage and what data loss cost the business”Risk if not established: Medium
Reliability is bought. Until someone knows the price of not having it, every architectural decision about redundancy is a guess wearing the clothes of engineering.
The figure you need is a cost per unit of downtime, and it is rarely just lost revenue. Contractual penalties, staff who cannot work, recovery labour, regulatory notification duties, and the cost of re-acquiring a customer who left all belong in it. For some workloads the honest answer is that an hour of downtime costs almost nothing, and that is a useful answer: it tells you to stop over-engineering.
Data loss needs its own number, because it behaves differently. Downtime cost accumulates with time; data loss cost is often a step function. Losing five minutes of transactions may be an inconvenience, while losing an hour may be unrecoverable. The shape of that curve determines whether synchronous replication is worth its latency.
The common failure is to skip this because the business conversation is harder than the technical one. What follows is a workload over-engineered where the team found the problem interesting and under-engineered everywhere else, with no way to tell which is which.
On STACKIT. Nothing on the platform produces this number for you. It comes from the business, and no cloud provider can supply it.
Two things the platform does help with. The Cost Dashboard
gives you the other side of the equation, which is what the reliability you are considering would
actually cost to run. And STACKIT publishes a
service certificate for every service, which
tells you what a given level of reliability costs in architecture rather than in guesswork. See
REL 1.3.
Tradeoffs. None between pillars. The cost is stakeholder time, mostly outside engineering, and the answer is frequently uncomfortable. There is no way to arrive at a defensible target without it.
Verify. For your highest-ranked flow, what does one hour of unavailability cost, what does one hour of lost data cost, and who produced those figures?
REL 1.2 Set availability, RTO and RPO per critical flow rather than per workload
Section titled “REL 1.2 Set availability, RTO and RPO per critical flow rather than per workload”Risk if not established: Medium
A workload-level target is wrong for at least one part of the workload. The checkout path and the monthly report have legitimately different requirements, and a single number applied to both either over-protects the report or under-protects the checkout.
Three separate figures are needed, and they are frequently confused:
- Availability is the fraction of time the flow works. It sets your redundancy.
- RTO, the recovery time objective, is how long the flow may stay broken. It sets your recovery procedure and how often you rehearse it.
- RPO, the recovery point objective, is how much recent data may be lost. It sets your backup and replication frequency.
All three are business inputs, decided before the architecture rather than derived from it. A target that was calculated from what the current design achieves is not a target; it is a description.
The ranking of flows comes from REL 2, and doing it once serves several pillars: reliability
investment, cost scrutiny under COST 7, and performance targets under PERF 1 all follow the
same order.
On STACKIT. Per-flow targets are the natural unit here, because STACKIT’s own commitments are
made per service rather than per workload. A flow that touches Kubernetes Engine, PostgreSQL Flex
and Object Storage inherits three separate figures, and only the flow-level view composes them.
See REL 1.3.
Tradeoffs. More targets to agree, track and review. Cost Optimization benefits: per-flow targets are what allow you to spend less on the flows that need less, which a single workload-level number never permits.
Verify. For each critical flow, what are the availability target, the RTO and the RPO, and which flows differ from each other?
REL 1.3 Check every target against the published availability of the services it depends on
Section titled “REL 1.3 Check every target against the published availability of the services it depends on”Risk if not established: Medium
A target above what your dependencies support is a wish. Since dependencies on a flow are usually serial, their availabilities multiply: four services each committed to 99.9% give you roughly 99.6% before your own code has failed once.
So the check runs in two directions. Compose the published figures for everything the flow touches and compare the result against your target. Where the composition falls short, you have three honest options: add redundancy to raise the ceiling, remove a dependency from the critical path, or lower the target. Deciding to do nothing is a fourth option, and it needs to be recorded as an accepted risk rather than left implicit.
Read what the numbers actually mean, not just their size. Availability definitions differ in ways that matter, and the exclusions are usually where the surprises live.
On STACKIT. STACKIT publishes a service certificate for every service, each stating an agreed availability. Every figure this calculation needs is therefore on one page and stated per service, which is what makes it a calculation rather than an estimate.
The Compute Engine certificate also shows how much the deployment shape changes the figure, which is the lever you actually control:
| Deployment | Agreed availability, calendar month average |
|---|---|
| Single VM in one availability zone | 99.5% |
| VM in a Metro Availability Zone | 99.8% |
| System group, two VMs in two availability zones in one region | 99.9% for at least one VM |
Reading any provider’s availability figure correctly means checking three things, and the STACKIT certificates state all three plainly rather than leaving them to be inferred.
What counts as unavailable. For Compute Engine it is the loss of external connectivity. Every
provider defines availability narrowly like this, because it has to be measurable without knowing
what your application does. It follows that your health model under REL 10 should be stricter
than the SLA definition: you want to detect the degradation that a connectivity check was never
designed to see.
Where one service ends and the next begins. Block Storage carries its own certificate, so a storage fault is measured against that document rather than against the VM’s. Compose both certificates when a flow depends on both. This is the normal shape of a service catalogue and the reason per-service certificates are useful in the first place.
What a zone boundary buys. Each availability zone has separate power, cooling and local
network connectivity. The certificate also states that several zones can be located in the same
building. Zone redundancy therefore addresses infrastructure failure, while separation across sites
is a different question, and regions are what answer it. The design consequence belongs in
REL 4.
Finally, check which side of the shared responsibility line each part of your RPO sits on. For
Compute Engine, backup and recovery are the customer’s responsibility, which is the usual division
for infrastructure services. STACKIT offers Server Backup
Management as a
managed option; the point is that your RPO should name which of the two it relies on rather than
assume. See REL 8.
Read the published figures per region rather than assuming one covers both eu01 and eu02, and
have a target that leans on the distinction name the region it was read for. Whether Metro
Availability Zones exist in both bears on the same decision and is covered under REL 4.4.
Tradeoffs. Cost Optimization. Raising the achievable ceiling means redundancy, and the cost does not scale linearly with the benefit: moving from a single VM to a two-zone system group buys an entire class of outage for roughly double the compute, while the next increment buys much less.
Verify. For your most critical flow, which STACKIT services does it depend on, what does each certificate commit to, and what does the composition of those figures allow compared with your stated target?
REL 1.4 Have each target agreed and recorded by someone accountable for the outcome
Section titled “REL 1.4 Have each target agreed and recorded by someone accountable for the outcome”Risk if not established: Medium
An unowned target does not survive contact with a budget. Under cost pressure, backup retention gets shortened, a standby gets downsized, a third replica is dropped. Each change is individually defensible, none is announced as a reduction in reliability, and the workload drifts away from what the business assumed without anyone deciding that it should.
A recorded target with a name against it converts each of those into a decision somebody has to make explicitly. That is its main purpose. Record what was agreed, who agreed it, when, and what the agreement was based on, because the reasoning ages faster than the number.
Set a trigger for revisiting rather than a calendar reminder nobody honours: a material change in
the business, a new regulatory obligation, or an architecture change that alters the dependency
composition from REL 1.3.
On STACKIT. No platform feature applies. This is an organizational practice, and the record belongs wherever your architecture decisions live rather than in the cloud environment.
Tradeoffs. Operational Excellence. Another artefact to keep current, and a stale target is worse than none because it carries false authority.
Verify. Who agreed the RTO for your most critical flow, on what date, and where is that recorded?
REL 1.5 Keep your internal objective stricter than any commitment you make to others
Section titled “REL 1.5 Keep your internal objective stricter than any commitment you make to others”Risk if not established: Medium
If your internal objective equals your external commitment, you have no margin. You discover you are in breach at the same moment your customer does, and every incident becomes a contractual event.
Set the internal objective tighter, and treat the gap as your working room. Crossing it is a
signal to act; crossing the external one is a failure. The size of the gap is a judgement about
how quickly you detect and recover, which means it should be informed by the RTO from REL 1.2
rather than picked round.
Keep the two labelled distinctly. An objective you set for yourself and a commitment someone can hold you to are different in kind, and conflating them either makes you over-cautious about internal goals or careless about external ones.
On STACKIT. The service certificates are STACKIT’s commitments to you, not objectives. Your commitments to your own customers sit on top of them and must absorb your own failure modes as well, which is why they can never simply restate the platform figure.
Tradeoffs. Little. A stricter internal objective may trigger work that a purely contractual reading would not require, which is the point of having one.
Verify. What is the difference between your internal availability objective and any commitment you have made externally, and what does that difference buy you in response time?
Related
Section titled “Related”REL 2Critical flows, which supplies the ranking these targets attach toREL 4Redundancy, where the composition fromREL 1.3turns into architectureREL 8Backup and restore, which the RPO drivesREL 9Disaster recovery, which the RTO drives- Reliability tradeoffs, particularly the conflict with Cost Optimization
- STACKIT service certificates