REL 10. How do you know the workload is healthy, and how do you test that it stays so?
Last updated on
Everything in this pillar so far is a claim: the workload tolerates a zone failure, the circuit breaker opens, the restore takes two hours. This question is where claims become evidence.
Two halves. A health model tells you the current state of each flow, expressed as a verdict
rather than a wall of metrics. Testing verifies that the responses decided in REL 3.2 do
what they were designed to do. Neither works without the other: a health model with nothing to
detect is decoration, and fault injection without a health model produces an experiment nobody can
read.
Best practices
Section titled “Best practices”REL 10.1Define what healthy means per flow, mapping signals to a verdictREL 10.2Alert on user-visible symptoms rather than on component metrics aloneREL 10.3Inject faults and verify the response matches the designREL 10.4Rehearse failover on a cadence, in production where you can
REL 10.1 Define what healthy means per flow, mapping signals to a verdict
Section titled “REL 10.1 Define what healthy means per flow, mapping signals to a verdict”Risk if not established: High
Monitoring tells you that a known condition occurred. A health model tells you whether a flow is working. The gap between those two is where most on-call time is spent: forty dashboards, six alerts firing, and nobody able to say whether customers are affected.
Build it per flow, using the map from REL 2.3. For each flow, define the states it can be in and
the signals that distinguish them. Three states minimum, because two are not enough:
- Healthy. Serving correctly within its targets.
- Degraded. Serving with reduced function or outside targets. This is the state
REL 6produces, and a two-state model cannot express it, which is why degradation goes unnoticed. - Failed. Not serving.
The mapping is the work. Which combination of component signals means the flow is degraded rather than failed, and which component failures do not affect this flow at all. That last part is what lets an operator ignore a firing alert with confidence, which is worth as much as the alerts themselves.
Include the failure modes from REL 3.1 that are hard to see. A component that is slow or
returning wrong answers passes most health checks, so the model needs signals that catch them:
latency percentiles rather than averages, and correctness checks rather than liveness.
On STACKIT. Observability provides metrics, logs and traces with dashboards and alerting, and is where the model is expressed once you have designed it. The alerting overview covers the mechanism.
The signals themselves come from your instrumentation, which is OPS 7. No platform can emit a
metric that says whether your checkout flow is working.
One point carried from REL 1.3: the SLA definition of unavailable is loss of external
connectivity, because a provider cannot measure whether your application is producing correct
answers. Your health model should therefore be stricter than the SLA definition by design, since
you want to detect the degradation a connectivity check was never meant to see.
Tradeoffs. Operational Excellence. A health model is a design artefact that ages with the architecture and needs maintaining. Cost Optimization. The signals it depends on carry telemetry cost, which is the retention tension described in the Operational Excellence tradeoffs.
Verify. For your most critical flow, which combination of signals means degraded rather than failed? Can someone on call answer “are customers affected” from one place, and how long does it take them?
REL 10.2 Alert on user-visible symptoms rather than on component metrics alone
Section titled “REL 10.2 Alert on user-visible symptoms rather than on component metrics alone”Risk if not established: High
Alerting on every component metric produces a volume nobody can act on, and the response to unactionable alerts is to stop reading them. Alert fatigue is a reliability problem, not a comfort problem: the alert that mattered arrives in a stream nobody trusts.
Alert on the symptom. Error rate and latency for a flow, measured where the user experiences them, catch every cause including the ones nobody predicted. Component alerts are for diagnosis after a symptom alert has fired, and for conditions that predict failure early enough to prevent it, such as a disk filling or a certificate expiring.
Two rules keep the volume honest. Every alert that pages a human needs a documented action, and if the action is “look at it and usually do nothing”, it is a dashboard rather than an alert. And every alert needs a stated urgency, because the difference between wake someone now and look at it Monday is the difference between a sustainable rotation and an exhausted one.
Review what fires. An alert that has never fired may be broken; one that fires weekly and is always dismissed is noise. Both are found by looking at the record rather than by intuition.
On STACKIT. Alert groups and
alerts
in Observability are where routing and severity are configured, and separating degraded from
failed here is what makes REL 6.4 operable.
status.stackit.cloud is the complementary signal for platform-side incidents. Checking it early during an investigation saves time that would otherwise go into looking for a cause in your own code.
Tradeoffs. Operational Excellence. Symptom-based alerting requires instrumentation at the
right layer, which is design work rather than configuration. Security. Alert routing is itself
on the recovery path, so REL 3.3 applies to it.
Verify. How many alerts fired in the last month, how many led to action, and how many pages were for conditions with no documented response?
REL 10.3 Inject faults and verify the response matches the design
Section titled “REL 10.3 Inject faults and verify the response matches the design”Risk if not established: High
REL 3.2 recorded a response for each failure mode. This is where you find out whether the
response happens. Untested reliability is assumed reliability, and the assumptions that fail do so
during real incidents.
Start from the failure mode list rather than from a tool. For each recorded response, design the smallest experiment that would falsify it: terminate an instance, block a dependency, fill a disk, add latency to a network path, exhaust a connection pool.
Follow the same shape every time. State the expected behaviour before running it, because an experiment without a prediction cannot fail. Limit the blast radius. Run it. Compare what happened with what you expected. The value is entirely in the difference.
Begin in a production-like environment and move to production where you can, since the
environments differ in exactly the ways that matter: scale, real traffic, real data volumes, real
dependency behaviour. OPS 6 is what makes the pre-production result meaningful at all.
Note what the exercise reveals about the health model. If a fault was injected and no signal
changed, REL 10.1 has a gap, and that finding is as valuable as the resilience result.
On STACKIT. No managed fault injection service exists, so the tooling is yours to choose and operate. For Kubernetes Engine the usual open-source chaos tooling runs in-cluster; for Compute Engine, terminating instances through the API or CLI is a legitimate and simple experiment that tests more than it appears to.
Zone-level experiments are worth designing carefully given the constraints in REL 4: in SKE, a
persistent volume is anchored to the zone where it was first bound, so terminating nodes in that
zone tests something quite different from terminating nodes elsewhere. That asymmetry is exactly
what an experiment should surface.
Tradeoffs. Reliability, in the short term: deliberately breaking things carries real risk, which is why blast radius limits and a tested rollback come first. Operational Excellence. Time and tooling for something that produces no feature.
Verify. Which failure modes from your analysis have been injected, when, and did the system behave as the design predicted? Which responses have never been tested?
REL 10.4 Rehearse failover on a cadence, in production where you can
Section titled “REL 10.4 Rehearse failover on a cadence, in production where you can”Risk if not established: High
Failover paths decay silently. Standby configuration drifts from primary, a certificate on the secondary path expires, replication falls behind, and a permission needed for promotion is removed during a cleanup. Nothing announces any of this.
The only reliable detection is regular exercise. Fail over deliberately, on a schedule, and treat it as routine rather than as an event. Teams that do this find that failover becomes boring, which is the goal: a mechanism exercised monthly is one people trust and can execute under pressure.
Rehearse in both directions. Failing back is frequently harder than failing over and is almost never practised, so the recovery ends with an environment running on its secondary indefinitely because nobody is confident about returning.
Time it, and compare with the RTO from REL 1.2, exactly as REL 9.3 does for full disaster
recovery. The difference is scope: this is component-level failover as routine maintenance, not a
declared disaster.
On STACKIT. What can be rehearsed depends on the service, and this is where REL 4.2 becomes
concrete: promotion behaviour for a PostgreSQL Flex replica set is what a rehearsal establishes
for your own configuration, the failover duration a recovery-time objective needs included.
For Kubernetes Engine, draining nodes in one zone is a well-bounded exercise that tests
scheduling, capacity headroom under REL 7.1, and the storage anchoring described in REL 4.2,
all at once.
Version updates are a rehearsal that happens anyway. SKE version updates and Compute Engine maintenance move workloads whether or not you planned an experiment, so treating them as scheduled resilience tests costs nothing and produces evidence.
Tradeoffs. Reliability. A rehearsal in production can cause a real incident, which is the honest cost. It is generally smaller than the cost of discovering the same defect during an unplanned outage, but it is not zero and it should be scheduled rather than surprising.
Verify. When did you last fail over each redundant component deliberately, how long did it take, and have you ever failed back?