---
title: Kubernetes Live Replication with Traffic Split
description: "Concrete template for Kubernetes-to-Kubernetes migration using live data replication and a staged traffic split to meet strict continuity requirements."
sidebar:
  badge:
    text: "STACKIT"
    variant: success
scfAsset:
  managed: false
  category: "runbook"
  external: false
  tags: ["design-and-mobilize", "use-cases", "replatform", "kubernetes", "replication", "traffic-split", "zero-downtime"]
  maintainers:
    - user: "lukas.weberruss"
      role: true
      website: true
source_url: "https://framework.stackit.cloud/migration/assetcontainer/stackit/runbook-k8s-live-replication-traffic-split/"
source_file: "docs/migration/assetcontainer/stackit/runbook-k8s-live-replication-traffic-split.mdx"
---

## Use case

- **Category**: Kubernetes-to-Kubernetes migration
- **State profile**: Stateful or mixed critical workloads
- **Approach**: Live replication plus staged traffic split

## Mandatory prerequisites

- **Replication path validated**: Data replication pipeline supports target RPO/RTO.
- **Traffic engineering available**: DNS, ingress, service mesh, or gateway supports weighted routing.
- **Dual-run operations model ready**: Ownership, monitoring, and incident playbooks exist for parallel runtime.
- **Rollback trigger policy approved**: Quantitative rollback thresholds are agreed before migration.

## Not suitable when

- **Replication lag cannot be controlled**: Freshness objectives cannot be met.
- **No progressive traffic control exists**: Cutover can only happen as full switch.

## Implementation template

### Phase 1: Establish dual-run baseline

1. Deploy target stack and validate health without production traffic.
2. Start replication pipeline and verify lag metrics.
3. Align alerting and SLO monitoring across source and target.

### Phase 2: Progressive traffic split

1. Shift low percentage traffic to target and observe error/latency budgets.
2. Increase traffic in controlled increments after each validation gate.
3. Keep source active until final confidence threshold is reached.

### Phase 3: Final switch and cleanup

1. Complete full traffic move to target.
2. Continue elevated monitoring through stabilization window.
3. Decommission source path after sign-off and evidence archive.

## Validation checklist

- **Replication health**: Lag and replay status within defined limits.
- **Traffic quality**: Error rate and latency stable at each split stage.
- **Business continuity**: Critical user journeys remain uninterrupted.
- **Rollback readiness**: Reverse split can run within agreed time.
