REL 9. How do you plan and rehearse disaster recovery?
Zuletzt aktualisiert am
Disaster recovery covers the failures that redundancy does not: the loss of a region, the destruction or corruption of the primary data set, a compromise that requires rebuilding from known-good state.
These events are rare enough that no team develops fluency through experience, which is exactly why the procedure has to be written and rehearsed. An unrehearsed plan is a set of assumptions, and the assumptions that turn out to be wrong are discovered under the worst conditions available.
Best practices
Section titled “Best practices”REL 9.1Define which disasters the plan covers, and which it does notREL 9.2Write the sequence, including who decides to invoke itREL 9.3Rehearse it and time the rehearsal against the RTOREL 9.4Keep the plan and its dependencies reachable when the primary environment is not
REL 9.1 Define which disasters the plan covers, and which it does not
Section titled “REL 9.1 Define which disasters the plan covers, and which it does not”Risk if not established: Medium
A plan for “a disaster” covers nothing specific. Recovery from a lost region, a corrupted database and a ransomware event share almost no steps, and a document that tries to cover all three at once will be too vague to follow.
Name the scenarios explicitly. The ones worth considering for most workloads:
- Loss of a region. Everything in one region is unavailable for an extended period.
- Loss or corruption of the primary data set. The data exists but is wrong, and the corruption may have replicated.
- Compromise. The environment cannot be trusted and must be rebuilt from known-good state.
- Loss of a critical dependency. A managed service or third party is unavailable beyond your tolerance.
Corruption deserves particular attention because redundancy actively works against you: a
corrupted write is replicated faithfully to every replica. Recovery is a restore to a point in
time, which makes it a backup problem under REL 8 rather than a redundancy problem.
State what is out of scope and why. A plan that quietly excludes the scenario the business assumed
was covered is worse than one that says plainly that a full region loss means a two-day outage,
accepted under REL 3.4.
On STACKIT. Two regions exist, eu01 in Germany and eu02 in Austria, described in
regions and availability zones . Both are inside EU
jurisdiction, which means a cross-region recovery design does not run into the sovereignty
constraint it might elsewhere. SOV 2 still requires the placement to be recorded, and the tier
from SOV 1 still governs.
status.stackit.cloud is the source for platform incident scope during an event, and its history is useful beforehand for calibrating which scenarios are worth planning against.
Tradeoffs. Operational Excellence. Analysis time, and the discomfort of naming what is not covered. Naming it is what turns an unexamined exposure into an accepted risk with an owner.
Verify. Which specific scenarios does your plan cover? For each scenario it does not cover, where is that exclusion recorded and who agreed to it?
REL 9.2 Write the sequence, including who decides to invoke it
Section titled “REL 9.2 Write the sequence, including who decides to invoke it”Risk if not established: High
Under pressure, people do not improvise well. The plan exists so that they do not have to.
Three parts are needed, and the first is the one most often missing.
The decision. Who is authorized to declare a disaster and invoke the plan, on what criteria, and who is called if that person is unreachable. Recovery is frequently delayed not by technical difficulty but by nobody being sure they were allowed to start. Invoking recovery is often irreversible, which is precisely why the authority needs to be settled in advance rather than negotiated during the event.
The sequence. Ordered steps, with dependencies between them made explicit. Written for someone competent who did not build the system, because the person who did may be unavailable. Name concrete resources rather than describing them.
The verification. How you confirm the recovery actually worked, before telling anyone it did. A restored system that is missing recent data or is subtly misconfigured is worse than an acknowledged outage.
Include communication. Who informs customers, who informs the regulator where that applies, and what is said while the outcome is still unknown.
On STACKIT. The platform provides the building blocks for recovery rather than a managed cross-region failover pattern, so the sequence, the data replication and the promotion are yours to design and operate. That is the usual division for infrastructure services, and planning for it explicitly is better than expecting a switch.
What the platform contributes is that the rebuild can be automated:
infrastructure defined as code through the Terraform provider, the
API or the CLI is what makes a recovery sequence executable rather than a list of console steps.
This is OPS 3 paying for itself, and it is the single largest factor in whether an RTO is
achievable.
One step of that sequence is worth naming now, because teams expect it to be free. Object Storage does not replicate between regions, so getting the objects into the second one is a job you write, schedule and monitor. Its interval is the recovery point for everything held there, which makes it a number in the plan rather than an implementation detail, and a job that silently stops is a recovery point quietly moving to whenever it last ran.
Tradeoffs. Operational Excellence. A plan that is not maintained becomes misleading, and
misleading during a disaster is expensive. Tie its review to the same trigger as the flow map in
REL 2.4.
Verify. Who is authorized to invoke your disaster recovery plan, and who is the deputy? Could someone who did not build the workload execute the sequence from the document as written?
REL 9.3 Rehearse it and time the rehearsal against the RTO
Section titled “REL 9.3 Rehearse it and time the rehearsal against the RTO”Risk if not established: High
A plan that has never been executed has unknown defects, and they are concentrated in the steps nobody thought about: the credential that expired, the DNS record with a long time to live, the dependency that was assumed to be present in the recovery environment.
Rehearse at a depth proportional to consequence. A tabletop walkthrough finds gaps in the decision
process and the sequence cheaply. A partial technical rehearsal restores one component into a
clean environment, which is REL 8.3. A full rehearsal recovers the workload into a separate
environment and verifies it serves traffic, which is the only version that produces a trustworthy
number.
Time every phase separately: detection, decision, execution, verification, and the return to normal operation. Teams are usually surprised by the decision phase rather than the execution phase.
Compare the total against the RTO from REL 1.2. If it exceeds the target, one of the two has to
change, and that is a conversation with whoever agreed the target under REL 1.4, not a problem
to absorb quietly.
Record what the rehearsal found. The value is in the defects discovered, and they need owners and
completion dates in the same way incident actions do under OPS 9.
On STACKIT. A rehearsal needs somewhere to recover into, and the ability to create and destroy
a full environment from code is what makes that affordable rather than a standing second copy. The
Terraform provider and CLI are the mechanism; the discipline is
OPS 3 and OPS 6.
For the data side, backup and
clone
in PostgreSQL Flex gives you a clean target without touching production, which is the property
REL 8.3 asks for and the same mechanism serves here.
Tradeoffs. Cost Optimization. A full rehearsal means standing up a second environment, even briefly, and it consumes engineering time that produces nothing visible. The alternative is an RTO nobody has evidence for.
Verify. When was the last full rehearsal, how long did it take end to end, and how did that compare with your stated RTO? What did it find, and were those items closed?
REL 9.4 Keep the plan and its dependencies reachable when the primary environment is not
Section titled “REL 9.4 Keep the plan and its dependencies reachable when the primary environment is not”Risk if not established: High
A recovery plan stored only in the environment being recovered is not available when needed. The same applies to everything the plan depends on.
Work through the circular dependencies deliberately, because they are invisible until tested:
- The plan document itself, if it lives in a wiki hosted on the affected infrastructure.
- Credentials and keys needed to authenticate to the recovery environment.
- Contact details for the people who must be reached, if the directory is unavailable.
- Artefacts: container images, packages, configuration. If the registry is inside the failure domain, there is nothing to deploy.
- The pipeline that deploys, if it runs in the affected environment.
Break-glass access belongs here too. An emergency path that bypasses normal controls is a
reliability necessity and a security liability at once, which is the tension named in the
Security tradeoffs. It needs to exist, be tightly bounded, be
heavily logged, and be reviewed after every use, per SEC 5.
On STACKIT. Authentication to the platform is on your recovery path: if the STACKIT Portal,
the API or the CLI is needed to rebuild, then identity is a recovery dependency, which is the
point REL 3.3 makes generally. Where identity is federated to an external provider, SOV 9 is
relevant to the same question from a different angle.
If your artefacts live in Container Registry or your source in STACKIT Git , both are components on the recovery path and belong in the analysis alongside everything else, not assumed to be present.
Tradeoffs. Security. Copies of credentials and plans held outside the primary environment are additional places they can leak from, and they need protection proportional to what they unlock.
Verify. If your primary region were unavailable right now, could you read the recovery plan, authenticate, and pull the artefacts needed to deploy? Which of those has been tested rather than assumed?
Related
Section titled “Related”REL 1Reliability targets, which supply the RTO the rehearsal is measured againstREL 3Failure modes, particularly the recovery dependencies inREL 3.3REL 8Backup and restore, which disaster recovery depends on entirelyOPS 8Operational procedures andOPS 9Incident managementSOV 2Placement and residency, which constrains where recovery may happen