---
id: REL10
pillar: reliability
title: REL 10. How do you know the workload is healthy, and how do you test that it stays so?
description: Monitoring produces alerts; a health model produces answers. How to map signals to a per-flow verdict and verify the design by injecting real faults into it.
status: draft
services: [observability, kubernetes-engine]
sidebar:
  order: 19
  label: Health model & testing
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/rel-10-health-model-and-testing/"
source_file: "docs/architecture/pillars/reliability/rel-10-health-model-and-testing.mdx"
---

Everything in this pillar so far is a claim: the workload tolerates a zone failure, the circuit
breaker opens, the restore takes two hours. This question is where claims become evidence.

Two halves. A **health model** tells you the current state of each flow, expressed as a verdict
rather than a wall of metrics. **Testing** verifies that the responses decided in [`REL 3.2`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-32-decide-and-record-the-response-to-each-failure-mode) do
what they were designed to do. Neither works without the other: a health model with nothing to
detect is decoration, and fault injection without a health model produces an experiment nobody can
read.

## Best practices

- [`REL 10.1`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-101-define-what-healthy-means-per-flow-mapping-signals-to-a-verdict) Define what healthy means per flow, mapping signals to a verdict
- [`REL 10.2`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-102-alert-on-user-visible-symptoms-rather-than-on-component-metrics-alone) Alert on user-visible symptoms rather than on component metrics alone
- [`REL 10.3`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-103-inject-faults-and-verify-the-response-matches-the-design) Inject faults and verify the response matches the design
- [`REL 10.4`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-104-rehearse-failover-on-a-cadence-in-production-where-you-can) Rehearse failover on a cadence, in production where you can

---

## REL 10.1 Define what healthy means per flow, mapping signals to a verdict

**Risk if not established:** High

Monitoring tells you that a known condition occurred. A health model tells you whether a flow is
working. The gap between those two is where most on-call time is spent: forty dashboards, six
alerts firing, and nobody able to say whether customers are affected.

Build it per flow, using the map from [`REL 2.3`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-23-map-each-flow-to-every-component-and-dependency-it-touches). For each flow, define the states it can be in and
the signals that distinguish them. Three states minimum, because two are not enough:

- **Healthy.** Serving correctly within its targets.
- **Degraded.** Serving with reduced function or outside targets. This is the state [`REL 6`](/architecture/pillars/reliability/rel-06-graceful-degradation/)
  produces, and a two-state model cannot express it, which is why degradation goes unnoticed.
- **Failed.** Not serving.

The mapping is the work. Which combination of component signals means the flow is degraded rather
than failed, and which component failures do not affect this flow at all. That last part is what
lets an operator ignore a firing alert with confidence, which is worth as much as the alerts
themselves.

Include the failure modes from [`REL 3.1`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-31-enumerate-how-each-component-on-a-critical-flow-can-fail) that are hard to see. A component that is slow or
returning wrong answers passes most health checks, so the model needs signals that catch them:
latency percentiles rather than averages, and correctness checks rather than liveness.

**On STACKIT.**
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip>
provides metrics, logs and traces with dashboards and alerting, and is where the model is
expressed once you have designed it. The <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/getting-started/alerting-overview/">alerting
overview</LinkChip>
covers the mechanism.

The signals themselves come from your instrumentation, which is [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/). No platform can emit a
metric that says whether your checkout flow is working.

One point carried from [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on): the SLA definition of unavailable is loss of external
connectivity, because a provider cannot measure whether your application is producing correct
answers. Your health model should therefore be stricter than the SLA definition by design, since
you want to detect the degradation a connectivity check was never meant to see.

**Tradeoffs.** **Operational Excellence.** A health model is a design artefact that ages with the
architecture and needs maintaining. **Cost Optimization.** The signals it depends on carry
telemetry cost, which is the retention tension described in the
[Operational Excellence tradeoffs](/architecture/pillars/operational-excellence/tradeoffs/).

**Verify.** For your most critical flow, which combination of signals means degraded rather than
failed? Can someone on call answer "are customers affected" from one place, and how long does it
take them?

---

## REL 10.2 Alert on user-visible symptoms rather than on component metrics alone

**Risk if not established:** High

Alerting on every component metric produces a volume nobody can act on, and the response to
unactionable alerts is to stop reading them. Alert fatigue is a reliability problem, not a comfort
problem: the alert that mattered arrives in a stream nobody trusts.

Alert on the symptom. Error rate and latency for a flow, measured where the user experiences them,
catch every cause including the ones nobody predicted. Component alerts are for diagnosis after a
symptom alert has fired, and for conditions that predict failure early enough to prevent it, such
as a disk filling or a certificate expiring.

Two rules keep the volume honest. Every alert that pages a human needs a documented action, and if
the action is "look at it and usually do nothing", it is a dashboard rather than an alert. And
every alert needs a stated urgency, because the difference between wake someone now and look at it
Monday is the difference between a sustainable rotation and an exhausted one.

Review what fires. An alert that has never fired may be broken; one that fires weekly and is
always dismissed is noise. Both are found by looking at the record rather than by intuition.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/manage-alert-groups-and-alerts/">Alert groups and
alerts</LinkChip>
in Observability are where routing and severity are configured, and separating degraded from
failed here is what makes [`REL 6.4`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-64-make-the-degraded-state-visible-to-operators-and-to-users) operable.

<LinkChip href="https://status.stackit.cloud">status.stackit.cloud</LinkChip> is the complementary signal for platform-side
incidents. Checking it early during an investigation saves time that would otherwise go into
looking for a cause in your own code.

**Tradeoffs.** **Operational Excellence.** Symptom-based alerting requires instrumentation at the
right layer, which is design work rather than configuration. **Security.** Alert routing is itself
on the recovery path, so [`REL 3.3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-33-include-the-failure-modes-of-your-dependencies-not-only-of-your-own-code) applies to it.

**Verify.** How many alerts fired in the last month, how many led to action, and how many pages
were for conditions with no documented response?

---

## REL 10.3 Inject faults and verify the response matches the design

**Risk if not established:** High

[`REL 3.2`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-32-decide-and-record-the-response-to-each-failure-mode) recorded a response for each failure mode. This is where you find out whether the
response happens. Untested reliability is assumed reliability, and the assumptions that fail do so
during real incidents.

Start from the failure mode list rather than from a tool. For each recorded response, design the
smallest experiment that would falsify it: terminate an instance, block a dependency, fill a disk,
add latency to a network path, exhaust a connection pool.

Follow the same shape every time. State the expected behaviour before running it, because an
experiment without a prediction cannot fail. Limit the blast radius. Run it. Compare what happened
with what you expected. The value is entirely in the difference.

Begin in a production-like environment and move to production where you can, since the
environments differ in exactly the ways that matter: scale, real traffic, real data volumes, real
dependency behaviour. [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) is what makes the pre-production result meaningful at all.

Note what the exercise reveals about the health model. If a fault was injected and no signal
changed, [`REL 10.1`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-101-define-what-healthy-means-per-flow-mapping-signals-to-a-verdict) has a gap, and that finding is as valuable as the resilience result.

**On STACKIT.** No managed fault injection service exists, so the tooling is yours to choose and
operate. For
<LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/">Kubernetes Engine</LinkChip> the usual
open-source chaos tooling runs in-cluster; for Compute Engine, terminating instances through the
API or CLI is a legitimate and simple experiment that tests more than it appears to.

Zone-level experiments are worth designing carefully given the constraints in [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/): in SKE, a
persistent volume is anchored to the zone where it was first bound, so terminating nodes in that
zone tests something quite different from terminating nodes elsewhere. That asymmetry is exactly
what an experiment should surface.

**Tradeoffs.** **Reliability**, in the short term: deliberately breaking things carries real risk,
which is why blast radius limits and a tested rollback come first. **Operational Excellence.**
Time and tooling for something that produces no feature.

**Verify.** Which failure modes from your analysis have been injected, when, and did the system
behave as the design predicted? Which responses have never been tested?

---

## REL 10.4 Rehearse failover on a cadence, in production where you can

**Risk if not established:** High

Failover paths decay silently. Standby configuration drifts from primary, a certificate on the
secondary path expires, replication falls behind, and a permission needed for promotion is removed
during a cleanup. Nothing announces any of this.

The only reliable detection is regular exercise. Fail over deliberately, on a schedule, and treat
it as routine rather than as an event. Teams that do this find that failover becomes boring, which
is the goal: a mechanism exercised monthly is one people trust and can execute under pressure.

Rehearse in both directions. Failing back is frequently harder than failing over and is almost
never practised, so the recovery ends with an environment running on its secondary indefinitely
because nobody is confident about returning.

Time it, and compare with the RTO from [`REL 1.2`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-12-set-availability-rto-and-rpo-per-critical-flow-rather-than-per-workload), exactly as [`REL 9.3`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-93-rehearse-it-and-time-the-rehearsal-against-the-rto) does for full disaster
recovery. The difference is scope: this is component-level failover as routine maintenance, not a
declared disaster.

**On STACKIT.** What can be rehearsed depends on the service, and this is where [`REL 4.2`](/architecture/pillars/reliability/rel-04-redundancy/#rel-42-make-state-redundant-and-know-where-each-data-set-is-anchored) becomes
concrete: promotion behaviour for a PostgreSQL Flex replica set is what a rehearsal establishes
for your own configuration, the failover duration a recovery-time objective needs included.

For Kubernetes Engine, draining nodes in one zone is a well-bounded exercise that tests
scheduling, capacity headroom under [`REL 7.1`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-71-size-for-measured-peaks-and-keep-headroom-for-the-peak-you-did-not-predict), and the storage anchoring described in [`REL 4.2`](/architecture/pillars/reliability/rel-04-redundancy/#rel-42-make-state-redundant-and-know-where-each-data-set-is-anchored),
all at once.

Version updates are a rehearsal that happens anyway. <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/operations/version-updates/">SKE version
updates</LinkChip>
and <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server/basics/maintenance/">Compute Engine
maintenance</LinkChip> move
workloads whether or not you planned an experiment, so treating them as scheduled resilience tests
costs nothing and produces evidence.

**Tradeoffs.** **Reliability.** A rehearsal in production can cause a real incident, which is the
honest cost. It is generally smaller than the cost of discovering the same defect during an
unplanned outage, but it is not zero and it should be scheduled rather than surprising.

**Verify.** When did you last fail over each redundant component deliberately, how long did it
take, and have you ever failed back?

---

## Related

- [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) Failure modes, which supplies what to test
- [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/) Redundancy, whose mechanisms this verifies
- [`REL 6`](/architecture/pillars/reliability/rel-06-graceful-degradation/) Graceful degradation, which needs a three-state health model to be visible
- [`REL 9`](/architecture/pillars/reliability/rel-09-disaster-recovery/) Disaster recovery, the larger-scope rehearsal
- [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) Observability and [`OPS 9`](/architecture/pillars/operational-excellence/ops-09-incident-management/) Incident management
