Skip to content
Beta

REL 2. How do you identify and rank the critical flows?

Last updated on

Treating every component as equally important is the most expensive mistake in this pillar. It spreads redundancy evenly across parts that need it and parts that do not, which costs more than protecting the important paths properly and protects them less.

The ranking produced here is used well beyond reliability. Cost scrutiny under COST 7 and performance targets under PERF 1 follow the same order, so the work is done once and paid for three times.

  • REL 2.1 Enumerate flows by what they deliver, not by the components that implement them
  • REL 2.2 Rank flows by the consequence of failure rather than by traffic volume
  • REL 2.3 Map each flow to every component and dependency it touches
  • REL 2.4 Keep the map current as the architecture changes

REL 2.1 Enumerate flows by what they deliver, not by the components that implement them

Section titled “REL 2.1 Enumerate flows by what they deliver, not by the components that implement them”

Risk if not established: Medium

A flow is a path through the workload that delivers something a person or a business process cares about. “Complete a checkout” is a flow. “The payment service” is not; it is a component that several flows happen to use.

The distinction matters because reliability is experienced along flows and engineered along components. If you plan in components, you end up asking whether the payment service is reliable enough, which has no answer without knowing which flows depend on it and what they need.

Write flows in the language of the people who use them. If a business stakeholder cannot recognize the list, it is a component inventory wearing different labels.

Include the flows nobody demonstrates: the nightly settlement job, the monthly regulatory export, the path a support agent uses to correct a mistaken order. These are frequently more consequential than the ones on the product roadmap and are almost always missing from the first draft.

On STACKIT. No platform feature enumerates your flows, because only you know what the workload is for.

Distributed tracing helps you check the list once you have drafted it, by showing which paths requests actually take. Observability provides the collection and query side; the instrumentation that makes traces meaningful is yours to add, which is OPS 7.

Tradeoffs. None between pillars. The cost is a workshop rather than engineering time. The main risk is producing a list that is technically accurate and unrecognizable to the business, which is worse than no list because it looks finished.

Verify. Show your flow list to someone who represents the users. How many entries do they recognize, and which flows do they name that are missing?


REL 2.2 Rank flows by the consequence of failure rather than by traffic volume

Section titled “REL 2.2 Rank flows by the consequence of failure rather than by traffic volume”

Risk if not established: Medium

Traffic volume is a poor proxy for importance. The busiest endpoint is often a health check or an asset request. The flow that matters may run twice a month and carry the entire quarter’s revenue.

Rank by what REL 1.1 produced: the cost of the flow being unavailable, and the cost of losing its data. Where a business figure is unavailable, rank by ordered comparison instead, which is easier to agree and almost as useful. Asking “if exactly one of these two had to stay up, which one” gets a decisive answer where asking for absolute numbers stalls.

Resist a ranking where everything is critical. If more than a handful of flows sit in the top tier, the exercise has not been done. The purpose is to enable saying no to redundancy somewhere, and a list without a bottom cannot do that.

Time matters too. A flow can be critical during business hours and irrelevant overnight, or critical for three days at quarter end and idle the rest of the year. Recording that shape lets you scale protection with it rather than paying for the peak continuously, which connects directly to SUS 3.

On STACKIT. Nothing on the platform supplies this ranking. It is a business judgement.

Once the ranking exists, it becomes an input to platform decisions: which flows justify a multi-zone topology under REL 4, which justify the higher availability tiers described in the service certificates , and which can safely run on a single instance.

Tradeoffs. None between pillars. The cost is political: the ranking is an organizational artefact as much as a technical one, and the owners of lower-ranked flows will notice. Doing it explicitly is still better than the alternative, which is ranking by whoever argues most persistently.

Verify. Which flow is ranked lowest, and what reduced protection does that ranking actually buy? If the answer is none, the ranking is not being used.


REL 2.3 Map each flow to every component and dependency it touches

Section titled “REL 2.3 Map each flow to every component and dependency it touches”

Risk if not established: Medium

A ranked flow is only actionable once you know what it runs on. The map from flow to components is what turns “checkout must stay up” into a list of things that must therefore be redundant.

Follow the flow past the boundaries of your own code. Managed services, identity providers, DNS, certificate issuance, outbound integrations, and the CI pipeline that deploys the fix all sit on the path in ways that only become obvious during an incident.

Two dependency types are missed most often. Control-plane dependencies are things you need in order to recover rather than in order to run, such as the ability to authenticate or to deploy. Shared dependencies are components that several flows touch, where a failure affects more than the flow you were looking at.

The map does not need to be a diagram. A table of flow against component is easier to keep current and easier to query, which matters more than how it looks.

On STACKIT. The Resource Manager hierarchy is the natural place for this to be visible: when projects follow ownership and function, the resources a flow depends on are largely the resources in its projects. A hierarchy that grew organically will not give you that, which is one of several reasons SEC 2 and COST 2 both ask for a deliberate structure.

Distributed tracing through Observability can confirm the map for paths that are exercised in production. It will not show you the dependencies that only appear during recovery.

Tradeoffs. Operational Excellence. A map that is not maintained becomes misleading, and misleading is worse than absent during an incident. See REL 2.4.

Verify. For your highest-ranked flow, list every component and external dependency on its path, including what you need in order to deploy a fix. Which of those were missing from the last version of the map?


REL 2.4 Keep the map current as the architecture changes

Section titled “REL 2.4 Keep the map current as the architecture changes”

Risk if not established: Medium

Flow maps decay faster than most documentation because they capture relationships rather than structure, and relationships change with every feature.

Tie the update to something that already happens rather than to a review cadence nobody honours. A new dependency added in code review, a new managed service provisioned, or a change to the resource hierarchy are all moments where the question “does this change a flow map” can be asked cheaply.

Prefer a map that is partly derived over one that is entirely written. Anything you can generate from infrastructure as code or from traces will stay closer to the truth than prose, which is another return on OPS 3.

On STACKIT. Infrastructure defined as code, whether through the Terraform provider or the CLI , gives you a queryable description of what exists. It does not tell you which flow a resource serves, so the mapping from resource to flow remains a human annotation. Labelling resources consistently is what makes that annotation survivable.

Tradeoffs. Operational Excellence. Real ongoing effort for a benefit that only appears during incidents and planning. It is among the first things dropped under delivery pressure, and the decay is silent.

Verify. When was the flow map last changed, and what change to the architecture triggered it? If the last update predates the last architectural change, the map is describing a system that no longer exists.


  • REL 1 Reliability targets, which the ranking here attaches to
  • REL 3 Failure modes, which is applied per component on these flows
  • REL 4 Redundancy, where the ranking decides what gets protected
  • COST 7 Flow-based optimization, which reuses this ranking
  • PERF 1 Performance targets, which are also set per flow