Use case
Section titled “Use case”Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.
This runbook supplies the approval and operational sequence around the reference repository’s
scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement
are separate operations. A successful Terraform apply is not migration acceptance.
Scope and assumptions
Section titled “Scope and assumptions”The reference migrates the public schema and validates public.album using row count and a
deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source
export. A real workload needs its own compatible export, schema assessment, business tests,
and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration,
not replication or zero-downtime cutover.
The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.
Roles and ownership
Section titled “Roles and ownership”| Role | Accountable decision or evidence |
|---|---|
| Migration lead | Window, checkpoints, go/no-go authority, rollback deadline, and incident coordination |
| Application owner | All writers identified, source freeze, business acceptance, and post-cutover write reconciliation |
| Platform engineer | Approved Terraform plan, SKE access, Gateway and DNS readiness, suspended reconcilers |
| Database owner | Consistent source evidence, rehearsal, protected backup, restore integrity, and rollback execution |
| Operations owner | Telemetry, incident routing, recovery ownership, retention, and stabilization exit |
Migration decision gates
Section titled “Migration decision gates”- Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
- Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
- Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
- Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
- Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.
These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.
Pre-migration checks
Section titled “Pre-migration checks”Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.
- Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
- Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
- Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
- Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
- Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
- Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
- Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.
Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.
With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.
Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Run plan (cutover window)
Section titled “Run plan (cutover window)”Phase 1: Prepare source and target
Section titled “Phase 1: Prepare source and target”- Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
- Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
- Run
bash scripts/validate_gateway.shand inspect workload rollout and PostgreSQL connectivity. - Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
- Set
enable_springboot_hpa = false,enable_load_generator = false, anddeploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
Phase 2: Freeze and rehearse
Section titled “Phase 2: Freeze and rehearse”- Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
- Export the final consistent dump and manifest through the approved source procedure.
- Execute rehearsal from the Replatform repository using the approved artifact directory:
python3 scripts/migrate_postgres.py rehearse \ --artifacts ../stackit-cmf-Rehost-springboot/artifacts \ --evidence .tmp/migration-run- Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
- Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
Phase 3: Cutover and release
Section titled “Phase 3: Cutover and release”- Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
- Execute the explicitly confirmed cutover:
python3 scripts/migrate_postgres.py cutover \ --artifacts ../stackit-cmf-Rehost-springboot/artifacts \ --evidence .tmp/migration-run \ --source-write-frozen --confirm-target springmusic- Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
- Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
- Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
- Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.
Validation checklist
Section titled “Validation checklist”Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.
| Gate | Required evidence |
|---|---|
| Data integrity | Source manifest checksum, expected count and fingerprint match the restored target; migration journal identifies the correct project and database |
| Runtime | Original replica count restored, rollout healthy, no unexplained restart or connection failures |
| Traffic | Gateway accepted, HTTPRoute references resolved, DNS matches the Gateway, and actual client requests reach the intended target |
| Business behavior | Migrated albums visible and approved user journeys pass; any write test has an agreed cleanup and reconciliation plan |
| Database transport | Application and migration connections require TLS; ACL admits only approved SKE egress or explicitly approved overrides |
| Observability | Both scrape jobs have actual up=1 samples, exporter reports database health, and dashboard values correspond to the workload |
| Recovery | Pre-cutover backup, checksum, original fingerprint, and journal retained privately and copied to protected durable storage |
| Configuration | Reviewed post-migration plan has no unexplained resource drift; source and target operating states are recorded |
Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.
Rollback criteria and steps
Section titled “Rollback criteria and steps”Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.
Rollback triggers
Section titled “Rollback triggers”Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.
Target-database rollback
Section titled “Target-database rollback”- Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
- Restore the protected pre-cutover target using the original evidence directory:
python3 scripts/migrate_postgres.py rollback \ --evidence .tmp/migration-run --confirm-target springmusic- Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
- Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
- Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.
Source failback and interrupted runs
Section titled “Source failback and interrupted runs”Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.
If cutover failed before a valid backup was recorded, inspect the journal and database with the
database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed
process, inspect remaining springmusic-migration-* pods and the stopped Deployment before
resuming. The local migration lock does not coordinate different execution hosts.
Handover to operations
Section titled “Handover to operations”Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.
Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.
Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.
Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.
Evidence log template
Section titled “Evidence log template”| Checkpoint | Owner | Timestamp | Result | Evidence reference |
|---|---|---|---|---|
| Target and client path ready | Platform engineer | YYYY-MM-DD HH:MM | Pass/Fail | Protected record |
| Source freeze and final export | Application and DB owners | YYYY-MM-DD HH:MM | Pass/Fail | Manifest and approval |
| Final rehearsal accepted | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Rehearsal journal |
| Pre-cutover backup proven | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Backup checksum and restore result |
| Cutover and integrity accepted | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Migration journal |
| Business and traffic accepted | Application owner | YYYY-MM-DD HH:MM | Pass/Fail | Test and routing evidence |
| Drift and telemetry reviewed | Platform engineer | YYYY-MM-DD HH:MM | Pass/Fail | Plan and metric evidence |
| Handover or rollback completed | Migration lead | YYYY-MM-DD HH:MM | Pass/Fail | Signed decision |
Asset historyAdded Sep 26, 2026LWUpdatedNo updates · 1 bar = 1 week i
- LWLukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwner
Lukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwnerActive 10 of the last 12 weeks · 47 updateswww.linkedin.com/in/lukas-weberruß-a360b081