Skip to content
Beta

Replatform to STACKIT: Spring Boot on SKE and PostgreSQL Flex

Last updated on

Stackit LogoStackit Logo
STACKIT

Replatform to STACKIT: Spring Boot on SKE and PostgreSQL Flex

Replatform Spring Boot and PostgreSQL to SKE and PostgreSQL Flex: prepare source and target, rehearse the data migration, cut over, validate, and stabilize.

LIFT

Replatform Strategy and Tools

Confirm the two deliberate substitutions: a VM service becomes a Kubernetes Deployment, and VM-local PostgreSQL becomes PostgreSQL Flex. Preserve the Spring Music JAR and business behavior; this is Replatform rather than VM Rehost or application Refactor.

Design and mobilizeDesignReplatform In 2 trails
R-strategy migration method Decision flow from discovery to production with the seven R-strategies: Relocate, Rehost, Replatform, Repurchase, Refactor, Retain, and Retire. R-strategy migration methodFrom discovery and path selection through the seven R-strategies to validation, transition, and production.DiscoveryDiscoveryAssess / prioritizeAssess / prioritizeDetermine migration pathDetermine migration pathValidationValidationTransitionTransitionProductionProductionRelocateRelocate(move VM)Define Landing ZoneDefine Landing ZoneUse migration toolsUse migration toolsAUTOMATEMANUALInstallInstallConfigConfigDeployDeployValidation & handoverRehostingRehosting(move application)Define Landing ZoneDefine Landing ZoneUse migration toolsUse migration toolsAUTOMATEMANUALInstallInstallConfigConfigDeployDeployReplatformingReplatforming(lift and reshape)Define Landing ZoneDefine Landing ZoneMap Target PlatformMap Target PlatformAdapt Platform StackAdapt Platform StackRepurchasingRepurchasing(replace, drop and shop)Purchase COTS/SaaS and licensingPurchase COTS/SaaS and licensingMigrate business processMigrate business processRefactoringRefactoring(re-architecting applications)Redesign application/ infrastructure architectureRedesign application/ infrastructure architectureApp code developmentApp code developmentFull ALM/SDLCFull ALM/SDLCIntegrationIntegrationRetain/moveRetain/movekeep for now or move laterRetire/decommissionRetire/decommissionLanding zone foundationLanding zone foundationShared platform base for all paths
R-strategy method placing Replatform between Rehost and Refactor

Replatform keeps core application behavior but changes selected platform components to gain operational or economic benefits. It sits between Rehost and Refactor in change intensity.

Comparing Replatform, Rehost, and Refactor

Section titled “Comparing Replatform, Rehost, and Refactor”
  • Rehost: Move workload location with minimal platform or code change. Example: Spring Boot stays on VM, only cloud target changes.
  • Replatform: Keep application behavior, but change selected platform layers. Example: Spring Boot runtime moves from VM to Kubernetes while core business logic remains unchanged.
  • Refactor: Change code structure or architecture significantly to unlock additional capabilities. Example: split monolith into services, redesign persistence model, and rework integration contracts.
  • Runtime platform swap: VM-based application hosting to Kubernetes.
  • Data platform swap: Self-managed database on VM to managed PaaS database service.
  • Ops capability swap: Host-centric monitoring and deployment model to managed platform-native operations.
  • Connectivity/control swap: Ingress, DNS, and service exposure model adapted to managed platform patterns.

These are Replatform changes as long as the core product behavior and major code paths remain mostly stable.

  • Operational bottlenecks can be reduced through managed platform capabilities.
  • Moderate change tolerance exists, but full redesign is out of scope.
  • Scalability and reliability goals require infrastructure-level improvements.
  • Cost optimization target can be reached with selective platform substitution.

Platform component selection

Identify which layers should change (for example runtime, database operations, integration controls).

Compatibility boundaries

Validate technical constraints and fallback options before introducing platform changes.

Risk-managed sequencing

Stage changes to avoid coupling too many unknowns in one cutover window.

Evidence and acceptance

Define measurable improvements for performance, resilience, and operational load.

  1. Define Landing Zone for the workload and its control boundaries.
  2. Map Target Platform components for runtime, data, and integration.
  3. Adapt Platform Stack prerequisites with sequencing and rollback checkpoints.
  4. Use migration tools to execute the transition through the shared migration path.
  5. Validate non-functional requirements and approve handover.

For stateful workloads, define source and target data platform responsibilities before runtime cutover.

  • Source data ownership: Clarify who owns dump/export run and consistency checks.
  • Target data ownership: Clarify who owns managed database provisioning, access controls, and backup baseline.
  • Migration sequencing: Separate schema/data move from runtime switch and validate each gate independently.
  • Temporary access controls: Plan temporary data migration access and explicit rollback/removal checkpoints.
  • Approved design decision record with scope, assumptions, and governance sign-off.
  • Validation evidence package for security, compliance, and operational readiness.
  • Strategy-specific migration runbook draft from the Design phase.
  • Handover package for Migration Factory Setup and wave planning.
  • Replatform decision matrix with selected substitutions.
  • Compatibility and constraint assessment.
  • Sequenced migration and rollback design.
  • Target-state operations model.
  • Benefit metrics and acceptance criteria.

For a runnable example of a platform swap from VM to Kubernetes with Spring Boot, use:

Asset title
Framework
Asset type

Use the asset for the runnable VM-to-Kubernetes and VM-to-managed-database implementation details.

Define Landing Zone

Define landing zone controls and guardrails as the start condition for the Replatform path. Confirm platform prerequisites for runtime, data, and integration layers so substitutions can be introduced without breaking governance or operability.

Map Target Platform

Define the target platform mapping for the Replatform path across runtime, data, and integration services. Make dependencies explicit, including identity, networking, and data responsibilities, so each change can be validated before cutover.

Adapt Platform Stack

Specify required platform prerequisite changes and sequencing for controlled transition. Define rollback guardrails, readiness checks, and run ownership so wave delivery stays predictable when multiple platform layers change together.

OPS

Target Runtime Architecture

This architecture maps the VM-based Spring Boot and PostgreSQL source to a Kubernetes runtime and managed database on STACKIT. The same application JAR is retained while provisioning, deployment, traffic management, data recovery, and operational responsibilities change.

The reference baseline uses one SKE worker and PostgreSQL Flex, with Envoy Gateway, STACKIT DNS, and Observability. It does not deploy the additional services or multi-zone topology shown in the optional extension pattern below.

  • Runtime standardization: replace a systemd-managed Java process with a reproducible Deployment and health checks.
  • Database operations: move PostgreSQL to a managed service without redesigning the application schema.
  • Controlled platform change: qualify rollout, scaling, network access, and recovery independently before production acceptance.
Source VMApproved dump + manifestApplication clientsSTACKIT Application ProjectSpring Music JAR + systemdSelf-managed PostgreSQLSTACKIT DNSSKE: single-worker referencePostgreSQL FlexObservability + GrafanaEnvoy Gateway + HTTPRoutesClusterIP ServiceSame JAR on Java 11Boot 2 adapter + PG exporterTemporary migration clientManaged ExternalDNSspringmusicspringmusic_rehearsal local SQL watch route hostnamesJDBC / TLSrehearse / prove backupapproved cutover / rollbackdatabase metrics / TLSpublish Gateway addressscrape via Gateway 9090 / 9187freeze / export / verifyprotected transfer via kubectlresolve hostnameHTTP baseline; HTTPS optional

An init container verifies the commit-pinned JAR checksum before Java starts. Application containers are replaceable: authoritative album data lives in PostgreSQL Flex, not in a pod filesystem or Kubernetes PersistentVolume. Kubernetes Secrets inject database credentials; an external Secret Manager integration is not implemented in this baseline.

The Flex ACL defaults to actual SKE egress CIDRs. Both application and migration client require encrypted database connections. The migration client uses an isolated rehearsal database and only replaces the application data after explicit approval and a verified pre-cutover backup. No source-VM database connection or temporary public Flex ACL is required for the dump-based path.

Terraform installs Envoy Gateway and then a local routing chart. The application Service is ClusterIP; Envoy supplies the public LoadBalancer. SKE-managed ExternalDNS publishes the HTTPRoute hostname from the Gateway address. This is Gateway API, not a legacy Ingress controller or a separately provisioned STACKIT Application Load Balancer service.

HTTP is the tested default. For HTTPS, supply a trusted TLS Secret and configure gateway_tls_secret_name according to the repository procedure; certificate issuance and renewal remain external responsibilities. The separate metrics listeners are public and unauthenticated in the reference and require protection before sensitive use.

Boot 2 Actuator binds to pod-local loopback; the metrics adapter exposes selected measurements. The PostgreSQL exporter and the SKE monitoring integration feed Observability. Terraform creates the Grafana folder and dashboard, but dashboard availability alone does not establish application health, scrape continuity, or working alert delivery.

The tested worker count, HTTP endpoint, and sample application are a functional baseline, not an HA production architecture. Select a supported SKE release and suitable zone capacity. Assess multiple workers, zone distribution, workload disruption budgets, replica safety, database availability, and the traffic layer as separate design decisions with failure tests.

Database rollback restores the pre-cutover target, while Flex managed backups serve service recovery. Neither automatically redirects users to the source VM. Define write ownership, traffic-switch authority, rollback deadline, retention, and recovery objectives before migration.

The following broader design illustrates possible additions, not resources created by the reference Terraform. Additional node pools, topology rules, persistent volumes, RabbitMQ, Object Storage, and Secret Manager need their own implementation, ownership, and validation. Use them only for a demonstrated workload requirement; do not infer HA from this diagram.

InternetApplication ProjectBackend ServicesAccessKubernetes (SKE)PostgreSQLRabbitMQObject StorageSecret ManagerObservabilityExternal LBDNSEntry LayerService LayerWorkload LayerPlatform LayerGateway APIExternalDNSK8s ServicePersistent StorageDeploymentHPANode Pool AZ-1Node Pool AZ-2Node AutoscalerPVPVPod APod BVMVM
  • Decouple runtime and data migration gates: validate database migration and runtime rollout independently.
  • Design secret delivery explicitly: the reference uses Kubernetes Secrets and protected Terraform state; add a reviewed external secret integration when required.
  • Standardize observability labels and dashboards: make cross-application operation and incident handling consistent.
  • Keep RabbitMQ optional and explicit: add it when asynchronous integration or buffering is required.
  • Qualify availability separately: a multi-zone design needs suitable worker capacity, placement rules, disruption budgets, and application and database failure testing; it is not enabled by the baseline.
Cloud Framework Replatform Spring Boot with Terraform Follow the executable provisioning, source-evidence, rehearsal, cutover, and rollback workflow for this architecture. Open page Code & registry github.com STACKIT CMF Replatform Spring Boot Kubernetes repository Open the repository
  1. Copy the example file: cp env.tfvars.example env.tfvars
  2. Set required identity/project values:
service_account_key_path = "/path/to/stackit-sa-key.json"
create_project = true
target_project_owner_email = "owner@sa.stackit.cloud"
parent_container_id = "cmf-parent-container-id"
ske_cluster_name = "rpltfk8s01"
observability_instance_name = "cmf-rpltf-observability"
dns_zone_name = "cmf-example.runs.onstackit.cloud"
dns_zone_display_name = "cmf-example"
  1. Enable the target architecture switches:
observability_enabled = true
create_observability_instance = true
dns_enabled = true
create_dns_zone = true
deploy_workload = true
enable_postgres_flex = true
enable_springboot_hpa = false
enable_load_generator = false
springboot_replicas = 1
deploy_postgres_migration_job = false
create_grafana_dashboard = true
  1. Optional CMF flag wrapper (flags.env):
setup_project=true
setup_observability=true
setup_database=true
setup_workload=true
setup_loadgen=false
setup_dns=true
  1. Apply:
Terminal window
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan

Expected result: springboot_url reaches the application through the Gateway, the application uses PostgreSQL Flex, and grafana_dashboard_url opens the managed dashboard. Provisioning does not import source data. Follow the separate rehearsal and cutover workflow after target validation; keep HPA disabled throughout migration.

Code & registry github.com Implemented topology and prerequisites Review the exact resource definitions and operational boundaries in the Spring Boot Replatform repository. Open the repository
LIVE

Reference Implementation

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
STEP

Prepare the Lab and Repositories

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
SAFE

Prepare Source Evidence

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
STEP

Configure Target Inputs

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
BASE

Provision the Target

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
SAFE

Verify Database Access Boundaries

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
STEP

Configure kubectl and Validate the Target

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
OPS

Verify the Telemetry Path

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
STEP

Rehearse and Approve

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository

Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.

This runbook supplies the approval and operational sequence around the reference repository’s scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement are separate operations. A successful Terraform apply is not migration acceptance.

The reference migrates the public schema and validates public.album using row count and a deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source export. A real workload needs its own compatible export, schema assessment, business tests, and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration, not replication or zero-downtime cutover.

The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.

  • Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
  • Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
  • Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
  • Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
  • Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.

These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.

Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.

  • Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
  • Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
  • Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
  • Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
  • Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
  • Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
  • Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.

Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.

From the STACKIT docsCreate and manage instances › ACLSource updated 24.09.2026 · copied 06.10.2026

With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.

Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
  2. Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
  3. Run bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.
  4. Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
  5. Set enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
  1. Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
  2. Export the final consistent dump and manifest through the approved source procedure.
  3. Execute rehearsal from the Replatform repository using the approved artifact directory:
Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run
  1. Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
  2. Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
  1. Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
  2. Execute the explicitly confirmed cutover:
Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic
  1. Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
  2. Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
  3. Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
  4. Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.

Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.

Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.

Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.

Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.

  1. Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
  2. Restore the protected pre-cutover target using the original evidence directory:
Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic
  1. Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
  2. Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
  3. Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.

Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.

If cutover failed before a valid backup was recorded, inspect the journal and database with the database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed process, inspect remaining springmusic-migration-* pods and the stopped Deployment before resuming. The local migration lock does not coordinate different execution hosts.

Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.

Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.

Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.

Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.

Code & registry github.com Executable migration and recovery workflow Use the reference repository for exact command prerequisites, evidence formats, safeguards, and validation coverage. Open the repository
LIVE

Execute the Cutover

Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.

This runbook supplies the approval and operational sequence around the reference repository’s scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement are separate operations. A successful Terraform apply is not migration acceptance.

The reference migrates the public schema and validates public.album using row count and a deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source export. A real workload needs its own compatible export, schema assessment, business tests, and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration, not replication or zero-downtime cutover.

The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.

  • Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
  • Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
  • Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
  • Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
  • Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.

These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.

Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.

  • Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
  • Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
  • Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
  • Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
  • Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
  • Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
  • Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.

Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.

From the STACKIT docsCreate and manage instances › ACLSource updated 24.09.2026 · copied 06.10.2026

With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.

Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
  2. Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
  3. Run bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.
  4. Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
  5. Set enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
  1. Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
  2. Export the final consistent dump and manifest through the approved source procedure.
  3. Execute rehearsal from the Replatform repository using the approved artifact directory:
Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run
  1. Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
  2. Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
  1. Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
  2. Execute the explicitly confirmed cutover:
Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic
  1. Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
  2. Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
  3. Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
  4. Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.

Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.

Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.

Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.

Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.

  1. Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
  2. Restore the protected pre-cutover target using the original evidence directory:
Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic
  1. Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
  2. Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
  3. Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.

Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.

If cutover failed before a valid backup was recorded, inspect the journal and database with the database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed process, inspect remaining springmusic-migration-* pods and the stopped Deployment before resuming. The local migration lock does not coordinate different execution hosts.

Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.

Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.

Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.

Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.

Code & registry github.com Executable migration and recovery workflow Use the reference repository for exact command prerequisites, evidence formats, safeguards, and validation coverage. Open the repository
SAFE

Validate or Roll Back

Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.

This runbook supplies the approval and operational sequence around the reference repository’s scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement are separate operations. A successful Terraform apply is not migration acceptance.

The reference migrates the public schema and validates public.album using row count and a deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source export. A real workload needs its own compatible export, schema assessment, business tests, and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration, not replication or zero-downtime cutover.

The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.

  • Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
  • Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
  • Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
  • Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
  • Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.

These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.

Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.

  • Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
  • Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
  • Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
  • Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
  • Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
  • Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
  • Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.

Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.

From the STACKIT docsCreate and manage instances › ACLSource updated 24.09.2026 · copied 06.10.2026

With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.

Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
  2. Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
  3. Run bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.
  4. Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
  5. Set enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
  1. Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
  2. Export the final consistent dump and manifest through the approved source procedure.
  3. Execute rehearsal from the Replatform repository using the approved artifact directory:
Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run
  1. Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
  2. Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
  1. Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
  2. Execute the explicitly confirmed cutover:
Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic
  1. Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
  2. Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
  3. Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
  4. Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.

Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.

Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.

Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.

Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.

  1. Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
  2. Restore the protected pre-cutover target using the original evidence directory:
Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic
  1. Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
  2. Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
  3. Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.

Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.

If cutover failed before a valid backup was recorded, inspect the journal and database with the database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed process, inspect remaining springmusic-migration-* pods and the stopped Deployment before resuming. The local migration lock does not coordinate different execution hosts.

Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.

Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.

Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.

Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.

Code & registry github.com Executable migration and recovery workflow Use the reference repository for exact command prerequisites, evidence formats, safeguards, and validation coverage. Open the repository

Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.

This runbook supplies the approval and operational sequence around the reference repository’s scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement are separate operations. A successful Terraform apply is not migration acceptance.

The reference migrates the public schema and validates public.album using row count and a deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source export. A real workload needs its own compatible export, schema assessment, business tests, and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration, not replication or zero-downtime cutover.

The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.

  • Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
  • Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
  • Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
  • Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
  • Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.

These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.

Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.

  • Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
  • Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
  • Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
  • Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
  • Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
  • Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
  • Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.

Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.

From the STACKIT docsCreate and manage instances › ACLSource updated 24.09.2026 · copied 06.10.2026

With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.

Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
  2. Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
  3. Run bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.
  4. Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
  5. Set enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
  1. Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
  2. Export the final consistent dump and manifest through the approved source procedure.
  3. Execute rehearsal from the Replatform repository using the approved artifact directory:
Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run
  1. Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
  2. Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
  1. Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
  2. Execute the explicitly confirmed cutover:
Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic
  1. Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
  2. Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
  3. Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
  4. Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.

Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.

Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.

Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.

Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.

  1. Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
  2. Restore the protected pre-cutover target using the original evidence directory:
Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic
  1. Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
  2. Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
  3. Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.

Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.

If cutover failed before a valid backup was recorded, inspect the journal and database with the database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed process, inspect remaining springmusic-migration-* pods and the stopped Deployment before resuming. The local migration lock does not coordinate different execution hosts.

Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.

Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.

Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.

Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.

Code & registry github.com Executable migration and recovery workflow Use the reference repository for exact command prerequisites, evidence formats, safeguards, and validation coverage. Open the repository
GOAL

Stabilize the Reference Implementation

Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.

This runbook supplies the approval and operational sequence around the reference repository’s scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement are separate operations. A successful Terraform apply is not migration acceptance.

The reference migrates the public schema and validates public.album using row count and a deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source export. A real workload needs its own compatible export, schema assessment, business tests, and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration, not replication or zero-downtime cutover.

The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.

  • Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
  • Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
  • Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
  • Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
  • Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.

These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.

Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.

  • Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
  • Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
  • Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
  • Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
  • Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
  • Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
  • Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.

Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.

From the STACKIT docsCreate and manage instances › ACLSource updated 24.09.2026 · copied 06.10.2026

With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.

Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
  2. Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
  3. Run bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.
  4. Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
  5. Set enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
  1. Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
  2. Export the final consistent dump and manifest through the approved source procedure.
  3. Execute rehearsal from the Replatform repository using the approved artifact directory:
Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run
  1. Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
  2. Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
  1. Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
  2. Execute the explicitly confirmed cutover:
Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic
  1. Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
  2. Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
  3. Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
  4. Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.

Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.

Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.

Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.

Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.

  1. Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
  2. Restore the protected pre-cutover target using the original evidence directory:
Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic
  1. Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
  2. Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
  3. Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.

Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.

If cutover failed before a valid backup was recorded, inspect the journal and database with the database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed process, inspect remaining springmusic-migration-* pods and the stopped Deployment before resuming. The local migration lock does not coordinate different execution hosts.

Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.

Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.

Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.

Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.

Code & registry github.com Executable migration and recovery workflow Use the reference repository for exact command prerequisites, evidence formats, safeguards, and validation coverage. Open the repository

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
OPS

Verified Results and Remaining Limits

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository

This asset applies the Migration Framework to a Spring Boot and PostgreSQL Replatform: the application moves from a VM service to STACKIT Kubernetes Engine (SKE), and its database moves from self-managed PostgreSQL to STACKIT PostgreSQL Flex. The business function and application JAR stay unchanged; the runtime and database operating models change.

The reference repository is the source of truth for Terraform, Helm charts, pinned artifacts, migration scripts, and validation. Use a reviewed revision containing the Gateway API and scripts/migrate_postgres.py workflows described here; an older revision with a direct database-import Job does not implement this procedure.

Code & registry github.com STACKIT Spring Boot Kubernetes Replatform repository Open the Terraform, Gateway API, PostgreSQL migration, and Observability implementation used throughout this asset. Open the repository
  • Runtime: the identical Spring Music Spring Boot 2.4.0 JAR runs on Java 11 in a Kubernetes Deployment instead of under systemd.
  • Data: PostgreSQL Flex supplies dedicated application and rehearsal databases; JDBC and migration clients require TLS.
  • Traffic: Envoy Gateway, Gateway API HTTPRoutes, and SKE-managed ExternalDNS replace the VM endpoint.
  • Operations: Kubernetes health and resource controls replace host-service management; managed telemetry covers cluster, application, and database signals.
  • Migration: source evidence, isolated rehearsal, an explicitly approved cutover, and verified database rollback remain separate from infrastructure provisioning.

Cloud Foundry, Object Storage, microservice decomposition, and application modernization are not part of this implementation. The old sample application demonstrates platform substitution, not a recommendation to deploy an unsupported application stack in production.

Review the architecture asset before choosing capacity and network controls. It separates the implemented topology from production extensions such as public HTTPS, highly available workers, and protected metrics.

Cloud Framework Spring Boot on SKE with PostgreSQL Flex and Gateway API Review the implemented topology, runtime and data boundaries, and separately qualified production extensions. Open page
  1. Confirm Replatform suitability, source compatibility, landing-zone readiness, and ownership.
  2. Prepare a trusted PostgreSQL dump and integrity manifest independently of target provisioning.
  3. Review and apply Terraform for SKE, PostgreSQL Flex, Gateway, DNS, workload, and Observability.
  4. Validate the target, then rehearse the final source dump in the isolated rehearsal database.
  5. Freeze source writes, approve downtime, and run the gated cutover with a protected target backup.
  6. Accept application and data evidence or restore the pre-cutover target; switch traffic through the approved operator procedure.
  7. Retain evidence through stabilization and use representative telemetry for later optimization.

Use an isolated Linux lab environment with Git, Terraform, kubectl, curl, jq, getent, Python 3.11 or newer, and the PostgreSQL server and client tools. The Rehost sample scripts also require runuser, sha256sum, a postgres OS account, and root privileges to create and validate a temporary local database. Run the sample commands in that prepared lab environment, not on a production database host. Keep the generated private artifacts readable by the operator running the migration; do not make them world-readable.

Obtain reviewed commit IDs for both repositories from the reference maintainer and export them as REHOST_REVISION and REPLATFORM_REVISION before continuing. The Replatform revision must contain the Gateway API and gated migration workflow. Do not assume the remote default branch already contains the locally tested implementation; if the approved revision is not available, stop and obtain it before attempting the walkthrough.

Replace both placeholders with the reviewed full 40-character commit IDs, then set them in the same Bash terminal:

Terminal window
export REHOST_REVISION="REPLACE_WITH_REVIEWED_REHOST_COMMIT_ID"
export REPLATFORM_REVISION="REPLACE_WITH_REVIEWED_REPLATFORM_COMMIT_ID"

Run the following complete block from an empty working directory. The if check rejects missing, empty, or malformed IDs, including unchanged placeholders. Each && runs the next command only if the previous one succeeded. Git checks whether the commits are available; the format check alone does not verify approval or repository contents.

Terminal window
if [[ ! ${REHOST_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ||
! ${REPLATFORM_REVISION:-} =~ ^[0-9a-fA-F]{40}$ ]]; then
printf '%s\n' "Set both revision variables to reviewed full 40-character commit IDs." >&2
false
else
umask 077 &&
git clone https://github.com/stackitcloud/stackit-cmf-Rehost-springboot.git &&
git clone https://github.com/stackitcloud/stackit-cmf-replatform-springboot-k8s.git &&
git -C stackit-cmf-Rehost-springboot checkout --detach "$REHOST_REVISION" &&
git -C stackit-cmf-replatform-springboot-k8s checkout --detach "$REPLATFORM_REVISION" &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/migrate_postgres.py &&
test -f stackit-cmf-replatform-springboot-k8s/scripts/validate_gateway.sh &&
cd stackit-cmf-replatform-springboot-k8s &&
printf '%s\n' "Workspace ready. Continue from this Replatform checkout." || {
printf '%s\n' "Preparation failed. Resolve the error before continuing." >&2
false
}
fi

Continue only after Workspace ready appears. On failure the terminal stays open; later preparation commands are skipped. Existing or partially cloned directories are not removed or overwritten: inspect them and preserve local changes before retrying in a new empty working directory. On success, the terminal is in the Replatform checkout for the next steps.

Use an approved STACKIT project, service account, DNS delegation, SKE capacity, and protected Terraform backend. Obtain credentials through the approved secret channel, never from this Trail. The following commands assume this directory layout and a reviewed configuration; they do not establish a validated greenfield production deployment.

Qualify a consistent source dump and its manifest before data enters the migration workflow.

The Rehost reference supplies scripts/create_source_dump.sh and scripts/validate_source_dump.sh for its reproducible sample. Run them from that repository. Its artifacts directory supplies source-postgresql.dump and source-postgresql.manifest to the Replatform workflow. The manifest records version 1, table=public.album, row_count, album_fingerprint, and dump_sha256.

For the reproducible sample only, run the following from the Replatform checkout in the prepared lab environment. The scripts create a temporary PostgreSQL instance from the versioned sample SQL, export it, then perform an independent test restore. Use a fresh artifact directory; do not overwrite evidence from a migration already in progress.

Terminal window
umask 077
pushd ../stackit-cmf-Rehost-springboot
bash scripts/create_source_dump.sh
bash scripts/validate_source_dump.sh
popd

Expect successful validation of eight rows and a matching fingerprint. These commands do not read a source VM. Keep using these exact artifacts for rehearsal and cutover.

For a real source, replace the sample generator with an approved export procedure: freeze all writers and derive the custom-format dump and manifest from the same consistent source snapshot. Do not mistake the generated eight-album sample for an export of an arbitrary running VM. Verify PostgreSQL compatibility, extensions, ownership, and schema dependencies before export.

Only trusted dumps may be restored because they execute SQL. This implementation migrates the public application schema and checks public.album; it deliberately excludes Flex-managed schemas. Other workloads require their own invariants and an adapted schema scope.

Code & registry github.com Spring Boot Rehost source and sample export Use the Rehost repository for the identical application artifact and reproducible PostgreSQL source-evidence tools. Open the repository

Provision the target from a reviewed plan with explicit project, capacity, access, and DNS inputs.

Use Terraform, kubectl, curl, jq, getent, and Python 3.11 or newer on Linux. Copy env.tfvars.example to env.tfvars and adapt the actual variables in that file. Keep the tested provider lock file and immutable image and JAR references. Confirm SKE version availability, node-pool capacity, project permissions, and DNS delegation before planning.

From the STACKIT docsLifecycle of Kubernetes Engine › Kubernetes end-of-life datesSource updated 24.08.2026 · copied 05.10.2026

Starting with Kubernetes v1.33, we remove minor versions on the patch day that precedes the upstream maintenance end-of-life (EOL) date. The following table below lists the upstream EOL date for each Kubernetes minor version and the corresponding expiration date in SKE:

Please refer to the official Kubernetes Release History for up-to-date announcements of new versions.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

Prefer an approved Application Landing Zone project. Set create_project = false and provide its project ID and service account key path. Project creation is an alternative requiring an approved parent container and permissions; it is not a replacement for landing-zone governance. Never put credentials, state, saved plans, or migration evidence in version control.

For a new checkout, create the private variable file without overwriting an existing one:

Terminal window
umask 077
test -e env.tfvars || cp env.tfvars.example env.tfvars
chmod 600 env.tfvars

Edit this file before planning: supply the approved project and service-account path, region, supported SKE version, available node-pool flavor and zone, and delegated DNS settings from the repository example. Review backend access and locking, quotas, costs, and the HTTP/public-metrics limitations. The next section’s feature flags are not a complete environment configuration.

The optional common wrapper maps setup_project, setup_observability, setup_database, setup_workload, setup_loadgen, and setup_dns to the repository’s Terraform switches. These select provisioning scope only: enabling the database does not authorize data replacement and never replaces the separate migration approval gate.

These values select the complete workload and database path. They supplement, rather than replace, the project, region, node-pool, and DNS values in the repository example.

deploy_workload = true
dns_enabled = true
enable_postgres_flex = true
postgres_flex_target_database = "springmusic"
postgres_flex_target_app_acl_cidrs = []
observability_enabled = true
create_observability_instance = true
create_grafana_dashboard = true
enable_springboot_hpa = false
enable_load_generator = false
deploy_postgres_migration_job = false

Keep HPA and load generation disabled during migration. Choose alert settings deliberately; the end-to-end test did not validate alert delivery. springboot_image selects the Java runtime, not an unrelated prebuilt application image. The init container downloads the commit-pinned Rehost JAR and verifies its SHA-256 before startup. Mirror immutable artifacts into approved artifact and image services for production.

Run from the Replatform repository and review the saved plan before applying:

Terminal window
umask 077
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan
bash scripts/validate_gateway.sh

Access control and temporary migration ACL extension

Section titled “Access control and temporary migration ACL extension”

With empty application ACL inputs, Terraform uses the SKE cluster’s actual egress CIDRs for PostgreSQL Flex. Explicit application or legacy ACL values override that default and must be reviewed. Do not permit 0.0.0.0/0.

The temporary PostgreSQL client runs inside SKE and receives the dump through kubectl. It does not connect directly to the source VM, and no source or workstation CIDR needs temporary Flex access. JDBC and database tools use sslmode=require; this requires encryption but does not provide the hostname verification of verify-full. Protect credentials in Kubernetes Secrets and the Terraform backend, and validate stronger certificate verification where required.

Restore the final dump into the isolated rehearsal database and require matching evidence less than 24 hours old.

After infrastructure apply completes, run an isolated restore into springmusic_rehearsal. Adjust the source directory to the approved artifacts; keep the evidence path private.

Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run

Rehearsal validates the manifest, dump checksum, target identity, row count, and fingerprint without replacing the application database. After the source write freeze, rehearse the final dump again. Cutover requires matching evidence from less than 24 hours ago. A successful rehearsal of an older or different dump is not approval for the final input.

VM PostgreSQL to PostgreSQL Flex migration option

Section titled “VM PostgreSQL to PostgreSQL Flex migration option”

The old deploy_postgres_migration_job = true path is disabled by validation. Use the gated workflow instead. Stop HPA, load generation, and all other target writers. Suspend Terraform and GitOps reconciliation while the script controls the Deployment replica count. Its local lock protects one checkout, not concurrent operators on different machines.

Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic

The source-write flag is an operator attestation, not an automatic source shutdown. Cutover scales the application to zero, saves and checksums the pre-cutover target, proves that backup by restoring it into the rehearsal database, and only then restores the source transactionally. It verifies data before restarting the original replica count. Failure leaves the application stopped for investigation. Preserve the evidence journal and backup; do not overwrite them to retry.

Compare the database evidence with the source manifest, check application behavior through the Gateway, and confirm both metrics jobs are healthy. Traffic switching and business acceptance remain operator-controlled steps; the script does not change the source application’s endpoint.

To restore the protected pre-cutover target database:

Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic

Rollback checks target identity and backup integrity, saves the current target separately, restores the original data, and verifies its fingerprint before restarting. Post-cutover writes are not merged; retain the pre-rollback dump for explicit reconciliation. This is target-database rollback, not automatic failback to the source VM.

Database metrics visibility in Observability

Section titled “Database metrics visibility in Observability”

Terraform manages the SCF Replatform Grafana folder and eight-panel dashboard against the existing Thanos datasource. Use grafana_dashboard_url to open it. Cluster CPU, cluster memory, running pods, application requests, and PostgreSQL health and pressure support acceptance and later optimization. No manual dashboard import is required.

Application metrics come from the pod-local Boot 2 Actuator through the metrics adapter on port 9090; the PostgreSQL exporter serves port 9187. Check both actual scrape results, not only dashboard rendering. Missing telemetry is an investigation trigger, never proof of zero load.

Export a short-lived kubeconfig with private permissions and inspect the default namespace:

Terminal window
umask 077
mkdir -p .tmp
terraform output -raw kubeconfig > .tmp/replatform.kubeconfig
export KUBECONFIG="$PWD/.tmp/replatform.kubeconfig"
kubectl get deploy,svc,pods -n springboot
kubectl get gateway,httproute -n springboot
kubectl rollout status deployment/springboot -n springboot
bash scripts/validate_gateway.sh

The Gateway validator checks acceptance, resolved route references, DNS, and the application response. Remove the local kubeconfig after use and obtain a fresh one when it expires. Do not treat successful rollout alone as data or business acceptance.

Keep migration rollback distinct from Flex service recovery and retain protected evidence outside ephemeral execution environments.

Flex retention is configured explicitly, with a 32-day default in this reference. Managed database backups and the migration pre-cutover dump serve different purposes. Rehearsing the latter does not prove managed-service restore, point-in-time recovery, or application disaster recovery. Assign recovery ownership and test the required service recovery path separately.

Retain protected evidence and backups outside an ephemeral dev container until the rollback window closes. After a killed migration process, inspect leftover springmusic-migration-* pods before resuming; do not use Terraform to restart an unverified target.

This asset demonstrates how to preserve application behavior while changing the runtime and database operating models:

  • Platform substitution: run the same Spring Boot JAR on SKE, connect it to PostgreSQL Flex over TLS, and expose it through Gateway API and DNS.
  • Controlled data migration: qualify a source dump and manifest, rehearse an isolated restore, and require explicit approval and a verified target backup before cutover. Restore the pre-cutover target when rollback is required.
  • Independent validation: check data integrity, application responses, Gateway and DNS readiness, and actual scrape results. Review the Terraform plan for unexplained drift rather than treating successful provisioning as migration acceptance.
  • Repeatable observability: manage the Observability integration and eight-panel Grafana dashboard through Terraform. Use application and database signals for acceptance, stabilization, and later optimization.

This evidence does not establish a complete greenfield replay, migration from a live production source, zero downtime, high availability, public Gateway TLS, interactive IDP login, alert delivery, or managed Flex recovery. The tested single-worker HTTP setup exposes unauthenticated metrics; resolve those production requirements before using sensitive data. Upgrade the sample application and validate an appropriate supported Kubernetes release as separate controlled changes.

Code & registry github.com Reference configuration, scripts, and validation evidence Use the repository README and versioned implementation for exact prerequisites, variables, commands, and supported recovery boundaries. Open the repository
GOAL

Stabilize and Optimize

Return to the Migration Framework's Optimize loop: collect representative operating evidence, identify the limiting layer, implement one controlled change, and validate reliability, performance, and cost before keeping it.

MigrateOptimizeOverview In 7 trails

Optimize starts when workloads run on STACKIT and real operating data is available. The module converts post-cutover observations into measurable improvements for performance, stability, and cost efficiency.

Optimize is not a one-time task. It is an iterative cycle that can overlap with early stabilization and post-cutover care.

Many right-sizing and tuning decisions are only reliable under real load patterns. After cutover, teams can use production telemetry to separate assumptions from actual behavior.

  1. Collect runtime evidence: utilization, latency, error rates, throughput, and cost drivers.
  2. Identify bottlenecks and waste patterns at workload, platform, and data layers.
  3. Prioritize actions by business impact, risk reduction, and FinOps effect.
  4. Implement tuning changes in controlled increments.
  5. Validate outcomes against SLO, reliability, and cost targets.
  6. Feed lessons learned into future migration waves and operating standards.
  • Rightsizing: Align compute, storage, and network capacity with actual demand profiles.
  • Performance tuning: Improve latency and throughput through configuration, scaling, and architecture adjustments.
  • Reliability hardening: Reduce incident frequency through better resilience, observability, and failure handling.
  • FinOps controls: Improve cost transparency, remove waste, and optimize run-rate efficiency.

Optimization decisions should be based on runtime evidence, not assumptions. For practical implementation, combine workload telemetry, alerting, and controlled infrastructure changes.

  • Managed observability baseline: Use STACKIT Observability to collect metrics, logs, and traces with Grafana, Prometheus, Thanos, Loki, and Tempo.
  • Detection logic: Define explicit thresholds and observation windows for low utilization and overload conditions.
  • Run path: Apply rightsizing through IaC changes (for example VM flavor changes) with rollback checkpoints.
  • Validation loop: Re-measure SLO, error rates, and run-cost after each tuning increment.
Filters

Within a group every tick widens the list. Groups narrow each other.

Framework

Status

Topics

Asset title
Framework
Asset type

For Replatform workloads on Kubernetes, optimization spans multiple layers and should be coordinated as one control loop.

  • Pod scaling: Use HPA to adapt replica count to workload pressure with explicit min/max limits.
  • Node pool scaling: Keep sufficient cluster headroom and tune machine type (flavor) for CPU/memory density requirements.
  • Ingress scaling: Re-evaluate load balancer service plan when ingress throughput or connection behavior becomes a bottleneck.
  • Storage rightsizing: Select storage classes based on performance requirements for persistent workloads.
  • Validation discipline: Re-check latency, error rate, and cost after every incremental tuning change.

Primary inputs

Cutover reports, incident trends, SLO measurements, telemetry baselines, and cost reports.

Optimization outputs

Prioritized improvement backlog, validated tuning changes, and updated runbook standards.

Governance outcome

Clear trade-off decisions between performance, resilience, and cost with documented ownership.

  • Optimize follows technical migration delivery in Migrate.
  • Optimize can run in parallel with early post-cutover care activities, while ownership for this care model is covered in the Run phase.
  • Deeper architectural redesign remains in Refactor.
OPS

Continue the Same Reference Implementation

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail
OPS

Right-size PostgreSQL Flex

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail
OPS

Right-size the Pod Budget

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail
STEP

Qualify Pod and Worker Scaling

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail
STEP

Optional: Configure and Inspect HPA

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail
SAFE

Validate the Optimization

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail
Trail historyAdded Oct 4, 2026LWUpdatedNo updates · 1 bar = 1 week i
Maintainers
LWLukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwnerActive 10 of the last 12 weeks · 47 updatesSTACKITwww.linkedin.com/in/lukas-weberruß-a360b081Contributed in STACKIT