Zum Inhalt springen
Beta

REL 3. How do you analyse failure modes and decide what happens when each occurs?

Zuletzt aktualisiert am

Asking “how do we prevent this” has diminishing returns and no natural stopping point. Asking “what happens when this fails, and is that acceptable” has an answer you can verify.

This question is where a design stops being optimistic. Its output is a list of failure modes with a decided response against each, including the ones where the decision is to accept the consequence. An unexamined failure mode is not absent; it is simply unbudgeted.

  • REL 3.1 Enumerate how each component on a critical flow can fail
  • REL 3.2 Decide and record the response to each failure mode
  • REL 3.3 Include the failure modes of your dependencies, not only of your own code
  • REL 3.4 Record accepted risks explicitly, with who accepted them

REL 3.1 Enumerate how each component on a critical flow can fail

Section titled “REL 3.1 Enumerate how each component on a critical flow can fail”

Risk if not established: Medium

Work along the flow map from REL 2.3, component by component, and ask what failure looks like for each. The useful discipline is to enumerate failure modes rather than failure causes, because causes are unbounded and modes are not.

The modes worth covering for almost any component:

  • It is gone. The instance, the zone, or the service is unreachable.
  • It is slow. Still responding, well outside its normal latency. Frequently worse than gone, because callers wait instead of failing over.
  • It is wrong. Responding promptly with incorrect data. The hardest to detect and the most damaging.
  • It is full. Out of disk, connections, quota, or file handles.
  • It is flapping. Alternating between healthy and unhealthy fast enough to defeat health checks.

Partial and grey failures deserve deliberate attention because designs quietly assume components are either working or not. A node that accepts connections and cannot reach storage satisfies most health checks while serving nothing.

On STACKIT. The service certificates tell you what the platform measures, which is a useful starting point and not the whole list. For Compute Engine, unavailability is defined as loss of external connectivity, so the “slow” and “wrong” modes sit outside it by design. That is how availability is defined industry-wide, since a provider cannot measure whether your application is producing correct answers. It does mean your own analysis has to cover those modes.

Planned events belong in the enumeration too. Compute Engine maintenance and SKE version updates are documented, predictable, and still cause a component to go away for a period. A design that only handles unplanned failure will be surprised by the planned kind.

Tradeoffs. Operational Excellence. Costs analysis time up front, and the list is never complete. Bound it by the flow ranking from REL 2.2 rather than by trying to be exhaustive.

Verify. For the top three components on your most critical flow, which of the five modes above have you considered, and which have a decided response?


REL 3.2 Decide and record the response to each failure mode

Section titled “REL 3.2 Decide and record the response to each failure mode”

Risk if not established: Medium

An enumerated failure mode with no assigned response is a list entry, not a mitigation. Each one needs a decision, and there are only a few kinds available:

  • Tolerate it. Redundancy absorbs the failure, REL 4.
  • Degrade around it. The flow continues with reduced function, REL 6.
  • Fail fast. Return an error quickly rather than waiting, REL 5.
  • Recover from it. The flow stops and is restored, REL 8 and REL 9.
  • Accept it. Nothing is done, deliberately, REL 3.4.

Record which one you chose and why. The reasoning ages faster than the decision, and a future reader who cannot reconstruct it will either preserve a mitigation that is no longer needed or remove one that still is.

Write the response so it is testable. “The system handles node failure” is not testable. “Pods reschedule to another zone within two minutes, verified by draining a node” is, which is what REL 10.3 will exercise.

On STACKIT. Which responses are available depends on the service, and the platform documents the mechanisms rather than prescribing the choice.

For example, a node failure in Kubernetes Engine is tolerated if the node pool spans zones and the workload is schedulable elsewhere, and it becomes a recovery event if persistent storage anchors the pod to the failed zone. The mechanism is documented; whether it fits your flow is your decision. See REL 4.2.

Tradeoffs. Operational Excellence. Every mitigation is a thing to operate, tune and test. A design that tolerates everything is usually one nobody understands well enough to run under pressure, which is a net loss of reliability.

Verify. Pick a failure mode from your analysis. Which of the five responses was chosen, where is the reasoning recorded, and when was the response last exercised?


REL 3.3 Include the failure modes of your dependencies, not only of your own code

Section titled “REL 3.3 Include the failure modes of your dependencies, not only of your own code”

Risk if not established: Medium

Teams analyse the components they built and treat the platform underneath as a constant. It is not. Managed services fail, regions have incidents, certificate issuance breaks, DNS goes stale, and identity providers become unreachable at exactly the moment you want to log in and fix something.

Three categories go missing most reliably:

Recovery dependencies. What you need in order to fix things: authentication, the deployment pipeline, access to the container registry, the runbook itself. If these sit inside the failure domain, your recovery plan has a circular dependency.

Quiet dependencies. DNS resolution, time synchronization, certificate validity, and outbound network access to third parties. Nobody lists them and everything stops without them.

Data dependencies. The upstream feed that populates a cache, the export another team consumes. Their failure may not show in your monitoring at all.

On STACKIT. Two sources help here. status.stackit.cloud reports current and past service incidents, which is the honest record of what has actually failed rather than what could. Reading its history for the services on your flow is a cheap way to calibrate the analysis.

The service certificates state what each service covers and where the next one begins, which is what makes dependency composition possible. Block Storage is certified separately from Compute Engine, so a storage fault is measured against the storage document. This per-service structure is what lets you enumerate dependency failure modes individually instead of treating the platform as one opaque thing.

Tradeoffs. Operational Excellence. Extends the analysis considerably. Bound it to the flows that carry consequence, and accept that the tail of unlikely dependency failures will remain unanalysed.

Verify. If your identity provider were unreachable right now, could you still deploy a fix to production? What is that answer based on, and when was it last tested?


REL 3.4 Record accepted risks explicitly, with who accepted them

Section titled “REL 3.4 Record accepted risks explicitly, with who accepted them”

Risk if not established: Medium

Some failure modes are not worth mitigating. That is a legitimate and frequently correct decision. It stops being legitimate when it is implicit, because an unrecorded acceptance is indistinguishable from an oversight, and after an incident it will be treated as one.

A recorded acceptance carries four things: the failure mode, the consequence if it occurs, why mitigation was not worth its cost, and who agreed. The last one matters most. It converts a technical omission into a business decision that somebody owns, which is the same mechanism REL 1.4 uses for targets.

Set a revisit trigger rather than a date. Acceptances are usually justified by conditions such as low transaction volume or a workload being non-customer-facing, and those conditions change without anyone revisiting the decision that rested on them.

On STACKIT. No platform feature applies. This belongs wherever your architecture decisions are recorded.

One platform-adjacent input is worth capturing in the acceptance: if the accepted risk depends on a published availability figure, note which service certificate version you relied on. Certificates are versioned and dated, and an acceptance based on a superseded figure should be re-examined rather than assumed to still hold.

Tradeoffs. Very little beyond the effort of writing it down. The main resistance is cultural, since recording an accepted risk makes it visible, which is precisely the point.

Verify. Name three failure modes your workload does not mitigate. For each, where is the acceptance recorded and who agreed to it? If none are recorded, they are not accepted risks; they are unexamined ones.


  • REL 2 Critical flows, which supplies the components to analyse
  • REL 4 Redundancy, REL 5 Resilient interactions, REL 6 Graceful degradation, which are the available responses
  • REL 10 Health model and testing, which verifies the responses actually work
  • OPS 9 Incident management, where unanalysed failure modes surface the expensive way