---
id: REL02
pillar: reliability
title: REL 2. How do you identify and rank the critical flows?
description: Reliability spend follows a ranking of what the business actually needs. How to enumerate the flows through a workload and order them by consequence of failure.
status: draft
services: [observability]
sidebar:
  order: 11
  label: Critical flows
source_url: "https://framework.stackit.cloud/architecture/pillars/reliability/rel-02-critical-flows/"
source_file: "docs/architecture/pillars/reliability/rel-02-critical-flows.mdx"
---

Treating every component as equally important is the most expensive mistake in this pillar. It
spreads redundancy evenly across parts that need it and parts that do not, which costs more than
protecting the important paths properly and protects them less.

The ranking produced here is used well beyond reliability. Cost scrutiny under [`COST 7`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/) and
performance targets under [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) follow the same order, so the work is done once and paid for
three times.

## Best practices

- [`REL 2.1`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-21-enumerate-flows-by-what-they-deliver-not-by-the-components-that-implement-them) Enumerate flows by what they deliver, not by the components that implement them
- [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume) Rank flows by the consequence of failure rather than by traffic volume
- [`REL 2.3`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-23-map-each-flow-to-every-component-and-dependency-it-touches) Map each flow to every component and dependency it touches
- [`REL 2.4`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-24-keep-the-map-current-as-the-architecture-changes) Keep the map current as the architecture changes

---

## REL 2.1 Enumerate flows by what they deliver, not by the components that implement them

**Risk if not established:** Medium

A flow is a path through the workload that delivers something a person or a business process cares
about. "Complete a checkout" is a flow. "The payment service" is not; it is a component that
several flows happen to use.

The distinction matters because reliability is experienced along flows and engineered along
components. If you plan in components, you end up asking whether the payment service is reliable
enough, which has no answer without knowing which flows depend on it and what they need.

Write flows in the language of the people who use them. If a business stakeholder cannot recognize
the list, it is a component inventory wearing different labels.

Include the flows nobody demonstrates: the nightly settlement job, the monthly regulatory export,
the path a support agent uses to correct a mistaken order. These are frequently more consequential
than the ones on the product roadmap and are almost always missing from the first draft.

**On STACKIT.** No platform feature enumerates your flows, because only you know what the workload
is for.

Distributed tracing helps you check the list once you have drafted it, by showing which paths
requests actually take.
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip>
provides the collection and query side; the instrumentation that makes traces meaningful is yours
to add, which is [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/).

**Tradeoffs.** None between pillars. The cost is a workshop rather than engineering time. The main
risk is producing a list that is technically accurate and unrecognizable to the business, which is
worse than no list because it looks finished.

**Verify.** Show your flow list to someone who represents the users. How many entries do they
recognize, and which flows do they name that are missing?

---

## REL 2.2 Rank flows by the consequence of failure rather than by traffic volume

**Risk if not established:** Medium

Traffic volume is a poor proxy for importance. The busiest endpoint is often a health check or an
asset request. The flow that matters may run twice a month and carry the entire quarter's revenue.

Rank by what [`REL 1.1`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-11-quantify-what-an-outage-and-what-data-loss-cost-the-business) produced: the cost of the flow being unavailable, and the cost of losing
its data. Where a business figure is unavailable, rank by ordered comparison instead, which is
easier to agree and almost as useful. Asking "if exactly one of these two had to stay up, which
one" gets a decisive answer where asking for absolute numbers stalls.

Resist a ranking where everything is critical. If more than a handful of flows sit in the top
tier, the exercise has not been done. The purpose is to enable saying no to redundancy somewhere,
and a list without a bottom cannot do that.

Time matters too. A flow can be critical during business hours and irrelevant overnight, or
critical for three days at quarter end and idle the rest of the year. Recording that shape lets
you scale protection with it rather than paying for the peak continuously, which connects directly
to [`SUS 3`](/architecture/pillars/sustainability/sus-03-demand-shaping/).

**On STACKIT.** Nothing on the platform supplies this ranking. It is a business judgement.

Once the ranking exists, it becomes an input to platform decisions: which flows justify a
multi-zone topology under [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/), which justify the higher availability tiers described in the
<LinkChip href="https://stackit.com/en/gtc/service-certificates">service certificates</LinkChip>, and which can safely run
on a single instance.

**Tradeoffs.** None between pillars. The cost is political: the ranking is an organizational
artefact as much as a technical one, and the owners of lower-ranked flows will notice. Doing it
explicitly is still better than the alternative, which is ranking by whoever argues most
persistently.

**Verify.** Which flow is ranked lowest, and what reduced protection does that ranking actually
buy? If the answer is none, the ranking is not being used.

---

## REL 2.3 Map each flow to every component and dependency it touches

**Risk if not established:** Medium

A ranked flow is only actionable once you know what it runs on. The map from flow to components is
what turns "checkout must stay up" into a list of things that must therefore be redundant.

Follow the flow past the boundaries of your own code. Managed services, identity providers, DNS,
certificate issuance, outbound integrations, and the CI pipeline that deploys the fix all sit on
the path in ways that only become obvious during an incident.

Two dependency types are missed most often. **Control-plane dependencies** are things you need in
order to recover rather than in order to run, such as the ability to authenticate or to deploy.
**Shared dependencies** are components that several flows touch, where a failure affects more than
the flow you were looking at.

The map does not need to be a diagram. A table of flow against component is easier to keep current
and easier to query, which matters more than how it looks.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/resource-manager/">Resource Manager</LinkChip>
hierarchy is the
natural place for this to be visible: when projects follow ownership and function, the resources a
flow depends on are largely the resources in its projects. A hierarchy that grew organically will
not give you that, which is one of several reasons [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) and [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) both ask for a deliberate
structure.

Distributed tracing through
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> can
confirm the map for paths that are exercised in production. It will not show you the dependencies
that only appear during recovery.

**Tradeoffs.** **Operational Excellence.** A map that is not maintained becomes misleading, and
misleading is worse than absent during an incident. See [`REL 2.4`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-24-keep-the-map-current-as-the-architecture-changes).

**Verify.** For your highest-ranked flow, list every component and external dependency on its
path, including what you need in order to deploy a fix. Which of those were missing from the last
version of the map?

---

## REL 2.4 Keep the map current as the architecture changes

**Risk if not established:** Medium

Flow maps decay faster than most documentation because they capture relationships rather than
structure, and relationships change with every feature.

Tie the update to something that already happens rather than to a review cadence nobody honours.
A new dependency added in code review, a new managed service provisioned, or a change to the
resource hierarchy are all moments where the question "does this change a flow map" can be asked
cheaply.

Prefer a map that is partly derived over one that is entirely written. Anything you can generate
from infrastructure as code or from traces will stay closer to the truth than prose, which is
another return on [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/).

**On STACKIT.** Infrastructure defined as code, whether through the
<LinkChip href="https://docs.stackit.cloud/">Terraform provider or the CLI</LinkChip>, gives you a queryable description of
what exists. It does not tell you which flow a resource serves, so the mapping from resource to
flow remains a human annotation. Labelling resources consistently is what makes that annotation
survivable.

**Tradeoffs.** **Operational Excellence.** Real ongoing effort for a benefit that only appears
during incidents and planning. It is among the first things dropped under delivery pressure, and
the decay is silent.

**Verify.** When was the flow map last changed, and what change to the architecture triggered it?
If the last update predates the last architectural change, the map is describing a system that no
longer exists.

---

## Related

- [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) Reliability targets, which the ranking here attaches to
- [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) Failure modes, which is applied per component on these flows
- [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/) Redundancy, where the ranking decides what gets protected
- [`COST 7`](/architecture/pillars/cost-optimization/cost-07-flow-based-optimization/) Flow-based optimization, which reuses this ranking
- [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) Performance targets, which are also set per flow
