---
id: OPS09
pillar: operational-excellence
title: OPS 9. How do you manage incidents and learn from them?
description: An incident is expensive information. Whether it was worth the price depends on what happens afterwards, and whether the review looks past the proximate cause.
status: draft
services: [observability]
sidebar:
  order: 18
  label: Incident management
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/ops-09-incident-management/"
source_file: "docs/architecture/pillars/operational-excellence/ops-09-incident-management.mdx"
---

Incidents are unavoidable. What varies is how long they last, how much they cost, and whether the
same one happens again.

The two halves of this question pull in different directions. During an incident, the goal is to
restore service, and understanding is secondary. After it, the goal is understanding, and speed no
longer matters. Teams that conflate the two either delay recovery to investigate or restore
service and never find out why it broke.

## Best practices

- [`OPS 9.1`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-91-define-severity-roles-and-escalation-before-you-need-them) Define severity, roles and escalation before you need them
- [`OPS 9.2`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-92-separate-restoring-service-from-investigating-the-cause) Separate restoring service from investigating the cause
- [`OPS 9.3`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-93-review-for-structural-causes-rather-than-proximate-ones) Review for structural causes rather than proximate ones
- [`OPS 9.4`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-94-track-review-actions-to-completion) Track review actions to completion

---

## OPS 9.1 Define severity, roles and escalation before you need them

**Risk if not established:** High

Improvising an incident process while an incident is running produces the two failure modes you
would expect: too many people involved and nobody deciding, or one person overwhelmed and nobody
informed.

**Severity** determines the response, so the levels need to be distinguishable in the first minute
by someone under stress. Base them on customer impact rather than on technical scope, since that
is what determines urgency and is usually easier to assess quickly.

**Roles** matter more than headcount. Someone coordinates and decides, which is a full-time job
and not compatible with also debugging. Someone communicates outward. Someone investigates. In a
small incident one person may hold several, and naming them still prevents the case where everyone
assumes another person is handling communications.

**Escalation** covers what happens when the first responder cannot resolve it, is unreachable, or
when the severity increases. Include the path to whoever can authorize a costly decision, since
that is the delay [`REL 9.2`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-92-write-the-sequence-including-who-decides-to-invoke-it) also identifies.

Write it down and make it findable without a search, which is [`OPS 8.4`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/#ops-84-keep-procedures-reachable-when-the-environment-is-not).

**On STACKIT.** <LinkChip href="https://status.stackit.cloud">status.stackit.cloud</LinkChip> is the first check when a
cause is not immediately apparent, since a platform incident changes both your diagnosis and your
communication.

Where the platform is involved, the STACKIT support path is part of your escalation and belongs in
the process with its expected response times, rather than being looked up during the incident.

**Tradeoffs.** Little beyond the effort of agreeing it. The main risk is a process heavy enough
that people avoid declaring incidents, which produces unmanaged incidents rather than fewer.

**Verify.** Who decides that an incident is severity one, who coordinates, and who talks to
customers? Could the person on call tonight answer that without looking it up?

---

## OPS 9.2 Separate restoring service from investigating the cause

**Risk if not established:** Medium

During an incident, restoring service is the objective. Understanding why is frequently
unnecessary for that, and pursuing it first extends the outage.

Roll back, fail over, disable the feature, shed load. Any of those can restore service without
knowing the cause, and all of them are faster than a diagnosis. This is what makes [`OPS 4.2`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-42-make-rollback-an-ordinary-operation-and-exercise-it) and
[`OPS 5.3`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-53-roll-back-automatically-when-a-gate-fails) valuable beyond the deployment they were built for.

The tension is real: some mitigations destroy the evidence needed later. Restarting a process
discards its memory state; scaling out hides a leak. Where that applies, capture first, then
mitigate. A copy of the logs, a heap dump, a snapshot of the metrics takes a minute and preserves
the investigation.

Record a timeline while it is happening rather than reconstructing it afterwards. Memory of an
incident is unreliable within hours, and the timeline is most of the review.

**On STACKIT.**
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> holds
the signals during and after, which is what makes the investigation possible once service is
restored. The retention decisions from [`OPS 7.4`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-74-set-retention-by-the-value-of-each-signal-rather-than-uniformly) determine how much of the evidence still exists
by the time anyone looks.

The <LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> answers "what changed" for
platform resources, recorded by default at organization, folder and project scope. Its 90-day
window in the Portal is usually ample for an incident review and is worth remembering for the
incident whose origin turns out to be older than that.

**Tradeoffs.** **Reliability.** Capturing evidence before mitigating costs minutes of outage. That
is usually the right trade for a significant incident and the wrong one for a trivial one, which
means it is a judgement the coordinator makes rather than a rule.

**Verify.** In your last significant incident, was service restored before the cause was
understood? What evidence was preserved, and was any lost to the mitigation?

---

## OPS 9.3 Review for structural causes rather than proximate ones

**Risk if not established:** Medium

Proximate causes are rarely interesting. The certificate expired, the disk filled, the
configuration was wrong. Each is true, each suggests a fix that prevents that exact recurrence,
and none of them explains why the system was able to fail that way.

The useful questions are structural. Why did nothing detect it earlier. Why did recovery take
longer than expected. What made this failure mode possible. What else shares that property. The
last one is where the value concentrates, because it converts one incident into a class of
prevented ones.

A review that concludes with a person's name has found the cheapest available answer and stopped.
The question underneath is why the system permitted a normal human error to have that consequence,
and that question has an actionable answer where blame does not. This is [`OPS 1.2`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-12-make-incident-review-blameless-in-practice-not-only-in-policy) being tested.

Review near misses too. They carry most of the learning and none of the cost, and they are only
reported when reporting is safe.

Keep the output readable by people who were not there. A review that only makes sense to
participants cannot inform anyone else, which is most of its potential value.

**On STACKIT.** No platform feature applies to the review itself.

One thing worth stating explicitly because the capability exists: the
<LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> records who performed each action. Its
purposes are investigation and compliance. Using it to identify someone to blame will end the
reporting of near misses immediately and permanently, which costs far more than any individual
incident.

**Tradeoffs.** **Cost Optimization.** A structural review takes hours rather than minutes, and the
actions it produces are larger than a targeted fix. That is the trade: fixing the class costs more
than fixing the instance.

**Verify.** Read your last three incident reviews. For each, was the identified cause proximate or
structural, and did the review ask what else shares that property?

---

## OPS 9.4 Track review actions to completion

**Risk if not established:** Medium

A review whose actions are not completed was theatre. This is the most common failure in the whole
question, and it is invisible because the review itself looks like the deliverable.

Each action needs an owner, a date and a place where it is visible alongside other work. Actions
that live only in the review document compete with nothing and therefore lose to everything.

Expect them to be unpopular. They arrive unplanned, they compete with committed work, and they
address a problem that has already been mitigated. That is exactly why they need the explicit
allocation from [`OPS 1.4`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-14-fund-operational-work-explicitly-rather-than-expecting-it-to-fit-in-the-gaps) rather than the hope that capacity appears.

Prioritize honestly. Not every action is worth doing, and a review that produces fifteen actions
will complete three. Ranking them and closing the rest as accepted risk under [`REL 3.4`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-34-record-accepted-risks-explicitly-with-who-accepted-them) is more
honest than a list that quietly ages.

Watch for repeats. The same action appearing after three incidents means it was never done, or
that the underlying cause was misidentified. Both are worth knowing.

**On STACKIT.** No platform feature applies. Actions belong in whatever backlog the team actually
works from, which is the only property that matters.

**Tradeoffs.** **Cost Optimization.** Incident actions displace planned work, and the displacement
is the mechanism by which the system improves. A team that never displaces planned work is a team
whose incidents will recur.

**Verify.** How many actions from incident reviews in the last six months are still open, and how
many were closed as done? Has the same action appeared in more than one review?

---

## Related

- [`OPS 1.2`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-12-make-incident-review-blameless-in-practice-not-only-in-policy) Blameless culture, which determines whether reviews surface anything useful
- [`OPS 1.4`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-14-fund-operational-work-explicitly-rather-than-expecting-it-to-fit-in-the-gaps) Funding operational work, without which actions are not completed
- [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) Observability, which supplies the evidence
- [`REL 3`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/) Failure modes, which reviews feed back into
- [`SEC 11`](/architecture/pillars/security/sec-11-detection-and-response/) Detection and response, the security-specific version of this process
