---
title: "Disaster Recovery and Business Continuity"
description: "Defining RTO and RPO for cloud workloads, classifying DR scenarios, and establishing a tested business continuity strategy on STACKIT — for regulated industries and near-KRITIS environments."
sidebar:
  order: 10
  label: "Disaster Recovery"
source_url: "https://framework.stackit.cloud/adoption/adapting-operating-models/disaster-recovery/"
source_file: "docs/adoption/adapting-operating-models/disaster-recovery.mdx"
---

## What is at stake

Disaster recovery is the discipline nobody practises — until they have to. Then every decision that could have been made earlier costs money and trust.

For most organisations, cloud migration is the moment when DR requirements are defined in writing for the first time. This is the opportunity: clearly, proportionately, and tested.

## RTO and RPO — the two fundamental questions

Every system has an implicit recovery time and an acceptable data loss threshold. Most organisations do not know these figures — and pay the price when an incident occurs.

**Recovery Time Objective (RTO):** How long may a system be unavailable after an outage? An RTO of 4 hours means: after at most 4 hours, the system must be functioning again.

**Recovery Point Objective (RPO):** How much data loss is acceptable? An RPO of 1 hour means: at most the last 60 minutes of data may be lost.

**The important rule:** The more aggressive the RTO and RPO, the higher the ongoing costs. An RTO of 15 minutes with an RPO of 5 minutes requires hot-standby environments. An RTO of 24 hours with an RPO of 12 hours can be met with daily backups and manual processes.

The right answer is: proportionate, not maximalist.

## Four tier classes for cloud workloads

![Dr Overview](./files/dr-overview.svg)

### Tier 1 — Mission Critical (RTO &lt; 1h, RPO &lt; 15min)

Systems whose failure immediately leads to revenue loss, regulatory consequences, or safety risks. Examples: payment systems, core banking systems, industrial production control.

### Tier 2 — Business Critical (RTO &lt; 4h, RPO &lt; 1h)

Systems that must be restored within a few hours, but whose outage is tolerable in the short term. Examples: ERP systems, CRM, internal portals.

### Tier 3 — Business Important (RTO &lt; 24h, RPO &lt; 4h)

Systems that can be restored within one working day. Examples: reporting systems, secondary databases, near-test systems.

### Tier 4 — Standard (RTO &lt; 72h, RPO &lt; 24h)

Systems without immediate business criticality. Examples: development environments, archive systems, secondary analytics tools.

## DR and compliance — what regulators expect

For regulated industries, DR is not a voluntary best practice but a regulatory requirement.

**BAIT (banking):** Requires documented contingency plans, regular tests, and failure scenarios for all material IT systems. Recovery targets must be approved by the management board.

**KRITIS (critical infrastructure):** Obliges KRITIS operators to business continuity management in accordance with BSI IT-Grundschutz or ISO 22301. DR tests are mandatory; results must be documented.

**GDPR:** Art. 32 requires the ability to restore the availability of personal data after an incident. An RPO of more than 24 hours for GDPR-relevant systems is difficult to justify.

**DORA (Digital Operational Resilience Act, from January 2025):** Obliges financial institutions to comprehensive ICT risk management with explicit BCM requirements and mandatory resilience tests.

**STACKIT as a partner:** STACKIT operates data centres in Germany with ISO 27001 and SOC 2 certification. Data residency in de-01/de-02 ensures that backup data does not leave Germany — a decisive advantage for regulated industries.

## DR tests — why they rarely happen and how to change that

The most common DR mistake is never testing the plan. DR tests are regularly postponed because the production environment is considered too critical to risk a simulated failover, no clear ownership for DR tests exists, and tests are seen as effort without direct benefit.

**The solution:** Anchor DR tests as a regular engineering ritual, not as an exceptional event.

Four test formats in ascending intensity:

**1. Tabletop exercise (annually, low effort):** The team works through a failure scenario together — on the whiteboard. No systems touched. Identifies gaps in processes and accountabilities.

**2. Component test (quarterly):** A specific backup or failover mechanism is tested in the staging environment. Duration: a few hours.

**3. Partial failover test (semi-annually):** A non-production-critical system is actually switched to the standby site. Measures actual RTO. Duration: half a day.

**4. Full DR exercise (annually):** Complete simulation of a data centre outage for all Tier 1 and Tier 2 systems. Requires planning and alignment with stakeholders. Delivers the real evidence for regulators.

## Responsibilities in DR operations

<CardGrid>
  <Card title="CIO / IT Lead">
    Approves RTO/RPO targets, acts as escalation authority in an emergency, signs off DR tests and
    reports results to the board and regulators.
  </Card>
  <Card title="Platform Team">
    Implements and maintains DR infrastructure, conducts component tests, documents recovery
    procedures.
  </Card>
  <Card title="Application owners">
    Classify their systems into tier classes, define application-specific recovery steps,
    participate in tabletop exercises.
  </Card>
  <Card title="STACKIT">
    Guarantees availability of cloud infrastructure in accordance with SLA, provides zone-to-zone
    replication, delivers audit logs for regulatory evidence.
  </Card>
</CardGrid>

## Implementation steps

<Steps>
1. **Inventory and classify all systems** — Tier assignment for each workload, documented and confirmed by the business unit.

2. **Have RTO/RPO targets approved by the board or CIO** — Not as an IT decision, but as a business decision with cost implications.

3. **Implement DR architecture per tier** — From simple backups in object storage to hot standby in the second STACKIT zone.

4. **Document recovery procedures** — As a runbook, not a concept. Concrete steps that can be executed in an emergency without further questions.

5. **Conduct the first tabletop exercise** — Within 30 days of go-live. Output: a list of open items with owner and deadline.

6. **Establish and maintain a DR test calendar** — Quarterly component tests, semi-annual partial failover, annual full exercise.

</Steps>
