OPS 9. How do you manage incidents and learn from them?
Zuletzt aktualisiert am
Incidents are unavoidable. What varies is how long they last, how much they cost, and whether the same one happens again.
The two halves of this question pull in different directions. During an incident, the goal is to restore service, and understanding is secondary. After it, the goal is understanding, and speed no longer matters. Teams that conflate the two either delay recovery to investigate or restore service and never find out why it broke.
Best practices
Section titled “Best practices”OPS 9.1Define severity, roles and escalation before you need themOPS 9.2Separate restoring service from investigating the causeOPS 9.3Review for structural causes rather than proximate onesOPS 9.4Track review actions to completion
OPS 9.1 Define severity, roles and escalation before you need them
Section titled “OPS 9.1 Define severity, roles and escalation before you need them”Risk if not established: High
Improvising an incident process while an incident is running produces the two failure modes you would expect: too many people involved and nobody deciding, or one person overwhelmed and nobody informed.
Severity determines the response, so the levels need to be distinguishable in the first minute by someone under stress. Base them on customer impact rather than on technical scope, since that is what determines urgency and is usually easier to assess quickly.
Roles matter more than headcount. Someone coordinates and decides, which is a full-time job and not compatible with also debugging. Someone communicates outward. Someone investigates. In a small incident one person may hold several, and naming them still prevents the case where everyone assumes another person is handling communications.
Escalation covers what happens when the first responder cannot resolve it, is unreachable, or
when the severity increases. Include the path to whoever can authorize a costly decision, since
that is the delay REL 9.2 also identifies.
Write it down and make it findable without a search, which is OPS 8.4.
On STACKIT. status.stackit.cloud is the first check when a cause is not immediately apparent, since a platform incident changes both your diagnosis and your communication.
Where the platform is involved, the STACKIT support path is part of your escalation and belongs in the process with its expected response times, rather than being looked up during the incident.
Tradeoffs. Little beyond the effort of agreeing it. The main risk is a process heavy enough that people avoid declaring incidents, which produces unmanaged incidents rather than fewer.
Verify. Who decides that an incident is severity one, who coordinates, and who talks to customers? Could the person on call tonight answer that without looking it up?
OPS 9.2 Separate restoring service from investigating the cause
Section titled “OPS 9.2 Separate restoring service from investigating the cause”Risk if not established: Medium
During an incident, restoring service is the objective. Understanding why is frequently unnecessary for that, and pursuing it first extends the outage.
Roll back, fail over, disable the feature, shed load. Any of those can restore service without
knowing the cause, and all of them are faster than a diagnosis. This is what makes OPS 4.2 and
OPS 5.3 valuable beyond the deployment they were built for.
The tension is real: some mitigations destroy the evidence needed later. Restarting a process discards its memory state; scaling out hides a leak. Where that applies, capture first, then mitigate. A copy of the logs, a heap dump, a snapshot of the metrics takes a minute and preserves the investigation.
Record a timeline while it is happening rather than reconstructing it afterwards. Memory of an incident is unreliable within hours, and the timeline is most of the review.
On STACKIT.
Observability holds
the signals during and after, which is what makes the investigation possible once service is
restored. The retention decisions from OPS 7.4 determine how much of the evidence still exists
by the time anyone looks.
The audit log answers “what changed” for platform resources, recorded by default at organization, folder and project scope. Its 90-day window in the Portal is usually ample for an incident review and is worth remembering for the incident whose origin turns out to be older than that.
Tradeoffs. Reliability. Capturing evidence before mitigating costs minutes of outage. That is usually the right trade for a significant incident and the wrong one for a trivial one, which means it is a judgement the coordinator makes rather than a rule.
Verify. In your last significant incident, was service restored before the cause was understood? What evidence was preserved, and was any lost to the mitigation?
OPS 9.3 Review for structural causes rather than proximate ones
Section titled “OPS 9.3 Review for structural causes rather than proximate ones”Risk if not established: Medium
Proximate causes are rarely interesting. The certificate expired, the disk filled, the configuration was wrong. Each is true, each suggests a fix that prevents that exact recurrence, and none of them explains why the system was able to fail that way.
The useful questions are structural. Why did nothing detect it earlier. Why did recovery take longer than expected. What made this failure mode possible. What else shares that property. The last one is where the value concentrates, because it converts one incident into a class of prevented ones.
A review that concludes with a person’s name has found the cheapest available answer and stopped.
The question underneath is why the system permitted a normal human error to have that consequence,
and that question has an actionable answer where blame does not. This is OPS 1.2 being tested.
Review near misses too. They carry most of the learning and none of the cost, and they are only reported when reporting is safe.
Keep the output readable by people who were not there. A review that only makes sense to participants cannot inform anyone else, which is most of its potential value.
On STACKIT. No platform feature applies to the review itself.
One thing worth stating explicitly because the capability exists: the audit log records who performed each action. Its purposes are investigation and compliance. Using it to identify someone to blame will end the reporting of near misses immediately and permanently, which costs far more than any individual incident.
Tradeoffs. Cost Optimization. A structural review takes hours rather than minutes, and the actions it produces are larger than a targeted fix. That is the trade: fixing the class costs more than fixing the instance.
Verify. Read your last three incident reviews. For each, was the identified cause proximate or structural, and did the review ask what else shares that property?
OPS 9.4 Track review actions to completion
Section titled “OPS 9.4 Track review actions to completion”Risk if not established: Medium
A review whose actions are not completed was theatre. This is the most common failure in the whole question, and it is invisible because the review itself looks like the deliverable.
Each action needs an owner, a date and a place where it is visible alongside other work. Actions that live only in the review document compete with nothing and therefore lose to everything.
Expect them to be unpopular. They arrive unplanned, they compete with committed work, and they
address a problem that has already been mitigated. That is exactly why they need the explicit
allocation from OPS 1.4 rather than the hope that capacity appears.
Prioritize honestly. Not every action is worth doing, and a review that produces fifteen actions
will complete three. Ranking them and closing the rest as accepted risk under REL 3.4 is more
honest than a list that quietly ages.
Watch for repeats. The same action appearing after three incidents means it was never done, or that the underlying cause was misidentified. Both are worth knowing.
On STACKIT. No platform feature applies. Actions belong in whatever backlog the team actually works from, which is the only property that matters.
Tradeoffs. Cost Optimization. Incident actions displace planned work, and the displacement is the mechanism by which the system improves. A team that never displaces planned work is a team whose incidents will recur.
Verify. How many actions from incident reviews in the last six months are still open, and how many were closed as done? Has the same action appeared in more than one review?
Related
Section titled “Related”OPS 1.2Blameless culture, which determines whether reviews surface anything usefulOPS 1.4Funding operational work, without which actions are not completedOPS 7Observability, which supplies the evidenceREL 3Failure modes, which reviews feed back intoSEC 11Detection and response, the security-specific version of this process