Compute footprint
CloudMent AI-Powered Cloud Advisor
CloudMent
Last updated on
Plan the Spring Boot platform change to SKE and PostgreSQL Flex: discovery, architecture, landing-zone readiness, migration gates, handoff, and optimization.
Place the platform change within the complete Migration Framework: assess the workload, design SKE and PostgreSQL Flex, prepare the landing zone, migrate with explicit data gates, and stabilize before optimization. The application JAR stays the same; the runtime and database operating models change.
Establish an initial workload baseline for the Spring Boot VM, PostgreSQL data, dependencies, and capacity. Qualify migration intent and identify the unknowns that detailed Discovery must resolve before a platform decision.
Rapid Discovery provides a fast, automated baseline of the current environment across on-premises and cloud landscapes. The focus is on quantifying the existing IT portfolio in a short time window, so teams can make early migration and commercial decisions with confidence.
At this stage, quantity and distribution matter more than deep application relationships.
Rapid Discovery builds an initial inventory of infrastructure and platform assets, including:
Compute footprint
Storage baseline
OS landscape
Kubernetes baseline
Database inventory
These metrics create the first fact-based view of migration scope.
The output of Rapid Discovery is a core input for:
This allows program stakeholders to align on financial direction and technical baseline before detailed planning starts.
Rapid Discovery is intentionally not a full application-level analysis. It does not include deep interviews with every application owner and does not aim to fully map all runtime dependencies.
That depth is covered in the subsequent Discovery phase, where infrastructure exports are enriched with targeted assessments and owner input to build a complete application picture.
Typical input sources include exports such as spreadsheets or similar inventory files from existing environments. This phase can be accelerated with AI-assisted tooling that extracts the required baseline metrics from uploaded datasets.
Rapid Discovery therefore acts as a prerequisite for structured cost indication and for shaping a realistic target environment strategy.
Use AI-assisted discovery assets when workload descriptions and inventory inputs need to be turned into first assessment and design artifacts for expert review.




The following diagram shows why Rapid Discovery is performed: raw source data is processed by tooling into a decision-ready baseline that supports early price indication and initial target sizing.
A robust Rapid Discovery typically follows a clear sequence:
The objective is not a perfect target architecture. The objective is a reliable starting point with enough accuracy for early decisions.
Result quality depends heavily on source quality. Typical issues include duplicates, outdated entries, inconsistent naming, and missing performance data.
Recommended practice for this phase:
This keeps cost indications traceable and allows focused refinement in the subsequent Discovery phase.
Rapid Discovery provides the volume baseline for early cost modeling. Captured assets are translated into STACKIT-relevant consumption dimensions, for example:
Combined with operating assumptions (runtime profile, availability targets, growth trajectory), this produces a solid first price indication and an initial TCO corridor.
At the end of Rapid Discovery, the following outputs should be available at a minimum:
Consolidated asset baseline
Quantities per technology domain are consolidated in one baseline.
Meaningful segmentation
Assets are segmented by criticality, environment, and modernization potential.
Traceable assumptions
Assumptions and identified data gaps are documented transparently.
Initial cost indication
Cost ranges and primary drivers are available for early planning.
Prioritized candidates
A prioritized list for deeper Discovery activities is available.
These outputs establish the working baseline for architecture, planning, and governance in the next Assess steps.
Common Rapid Discovery risks include over-simplified categorization, incomplete source systems, or overestimating data maturity.
Proven countermeasures:
This keeps the phase fast while preserving decision quality.
The handover point is reached when quantities, technology classes, and primary cost levers are sufficiently visible and open questions are clearly documented.
In Discovery, these open items are addressed through targeted owner interviews, deeper assessments, and dependency/compliance/operations analysis to build the full application-level picture.
Confirm Java and PostgreSQL compatibility, state handling, scheduled writers, schema dependencies, data volume and change rate, downtime tolerance, recovery objectives, and representative demand. Validate those inputs with the application and database owners before selecting the target.
Discovery is one of the first and most critical modules in the Design and Mobilize phase. It refines Rapid Discovery results and adds the depth needed to make architecture and migration-wave decisions with confidence.
The primary objective is to establish a realistic, evidence-based understanding of the current IT landscape, business priorities, and organizational readiness before detailed target design and migration planning are finalized.
Complete baseline
Create a reliable application and infrastructure baseline that goes beyond pure quantities.
Dependency transparency
Identify technical and process dependencies to avoid hidden migration blockers.
Business alignment
Link technical findings with business criticality, timelines, and risk tolerance.
Planning readiness
Produce decision-ready input for target design and migration-wave planning.
Inventory
Comprehensive capture of servers, virtual machines, databases, middleware, and applications.
Dependency analysis
Mapping of communication paths and runtime dependencies between systems and applications.
Resource utilization
Analysis of actual CPU, memory, storage, and I/O behavior over a representative period.
Operational context
Collection of backup, patching, SLA, compliance, and operational constraints.
Application owner input
Structured questionnaires and interviews to validate assumptions and close data gaps.
In practice, Discovery is often run together with STACKIT partners. Partners typically use their own tooling landscape to collect and normalize technical data into a central repository. Many programs also trigger targeted questionnaires for application owners directly from these tools to enrich technical findings with business and operational context.
This combined model improves speed and consistency while keeping stakeholder validation built into the process.
Discovery intentionally combines two evidence streams that complement each other:
Neither stream is sufficient on its own. Technical evidence without owner context can misclassify critical workloads, while human input without technical grounding can hide coupling and capacity risks. Discovery quality depends on reconciling both streams into one decision-ready view.
The following diagram shows how Discovery transforms technical and stakeholder input into decision-ready outputs for the downstream modules.
During Discovery, tooling commonly applies the following analysis patterns:
These analyses establish the technical fact base. The human-driven stream then validates, prioritizes, and contextualizes these findings for executable migration decisions.
Use AI-assisted discovery assets to structure workload inputs, service mapping, readiness findings, and R-strategy signals before architects validate the resulting discovery baseline.




Discovery outputs are directly reused by the next modules in Design and Mobilize:
Design
Uses dependency, capacity, and risk insights to shape target architecture options.
Security and Compliance
Uses data classification and control gaps to define prioritized security requirements.
Landing Zone
Uses platform and governance constraints to define foundational setup decisions.
Migration Plan
Uses move groups, criticality, and sequencing constraints for realistic wave planning.
Operating Model and Business Case
Uses ownership, process impact, and value/risk signals for staffing and investment priorities.
At minimum, Discovery should produce the following outputs:
These outputs are essential prerequisites for continuing with detailed design work and a credible migration plan.
Confirm the two deliberate substitutions: a VM service becomes a Kubernetes Deployment, and VM-local PostgreSQL becomes PostgreSQL Flex. Preserve the Spring Music JAR and business behavior; this is Replatform rather than VM Rehost or application Refactor.
Replatform keeps core application behavior but changes selected platform components to gain operational or economic benefits. It sits between Rehost and Refactor in change intensity.
These are Replatform changes as long as the core product behavior and major code paths remain mostly stable.
Platform component selection
Identify which layers should change (for example runtime, database operations, integration controls).
Compatibility boundaries
Validate technical constraints and fallback options before introducing platform changes.
Risk-managed sequencing
Stage changes to avoid coupling too many unknowns in one cutover window.
Evidence and acceptance
Define measurable improvements for performance, resilience, and operational load.
For stateful workloads, define source and target data platform responsibilities before runtime cutover.
For a runnable example of a platform swap from VM to Kubernetes with Spring Boot, use:
Use the asset for the runnable VM-to-Kubernetes and VM-to-managed-database implementation details.
Define landing zone controls and guardrails as the start condition for the Replatform path. Confirm platform prerequisites for runtime, data, and integration layers so substitutions can be introduced without breaking governance or operability.
Define the target platform mapping for the Replatform path across runtime, data, and integration services. Make dependencies explicit, including identity, networking, and data responsibilities, so each change can be validated before cutover.
Specify required platform prerequisite changes and sequencing for controlled transition. Define rollback guardrails, readiness checks, and run ownership so wave delivery stays predictable when multiple platform layers change together.
Use versioned Terraform for infrastructure and Kubernetes resources, Helm for Envoy Gateway and routes, and a separate approved script for data migration. Keep provisioning and data replacement independently reviewable; this target does not require Ansible host configuration.
Automation ensures that landing-zone capabilities are reproducible, versioned, and tested instead of manually configured.
For migration landing zones, automation is the delivery backbone that connects platform APIs, IaC tools, developer workflows, and release controls into one reliable operating model.
Terraform or OpenTofu and Ansible solve different parts of one delivery workflow. Keep the boundary explicit so infrastructure changes remain reviewable and host configuration remains repeatable.
Terraform / OpenTofu
Own the infrastructure lifecycle: projects, networks, security controls, compute, storage, managed services, and the outputs required by configuration management.
Ansible
Own configuration inside the reachable target: operating-system packages, middleware, application artifacts, service units, and workload-level validation.
Do not use provisioners or ad hoc scripts to blur ownership between both layers. Triggering Ansible from Terraform can be a practical bridge, but each tool must remain independently understandable, testable, and rerunnable.
STACKIT
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
This architecture maps the VM-based Spring Boot and PostgreSQL source to a Kubernetes runtime and managed database on STACKIT. The same application JAR is retained while provisioning, deployment, traffic management, data recovery, and operational responsibilities change.
The reference baseline uses one SKE worker and PostgreSQL Flex, with Envoy Gateway, STACKIT DNS, and Observability. It does not deploy the additional services or multi-zone topology shown in the optional extension pattern below.
An init container verifies the commit-pinned JAR checksum before Java starts. Application containers are replaceable: authoritative album data lives in PostgreSQL Flex, not in a pod filesystem or Kubernetes PersistentVolume. Kubernetes Secrets inject database credentials; an external Secret Manager integration is not implemented in this baseline.
The Flex ACL defaults to actual SKE egress CIDRs. Both application and migration client require encrypted database connections. The migration client uses an isolated rehearsal database and only replaces the application data after explicit approval and a verified pre-cutover backup. No source-VM database connection or temporary public Flex ACL is required for the dump-based path.
Terraform installs Envoy Gateway and then a local routing chart. The application Service is ClusterIP; Envoy supplies the public LoadBalancer. SKE-managed ExternalDNS publishes the HTTPRoute hostname from the Gateway address. This is Gateway API, not a legacy Ingress controller or a separately provisioned STACKIT Application Load Balancer service.
HTTP is the tested default. For HTTPS, supply a trusted TLS Secret and configure
gateway_tls_secret_name according to the repository procedure; certificate issuance and
renewal remain external responsibilities. The separate metrics listeners are public and
unauthenticated in the reference and require protection before sensitive use.
Boot 2 Actuator binds to pod-local loopback; the metrics adapter exposes selected measurements. The PostgreSQL exporter and the SKE monitoring integration feed Observability. Terraform creates the Grafana folder and dashboard, but dashboard availability alone does not establish application health, scrape continuity, or working alert delivery.
The tested worker count, HTTP endpoint, and sample application are a functional baseline, not an HA production architecture. Select a supported SKE release and suitable zone capacity. Assess multiple workers, zone distribution, workload disruption budgets, replica safety, database availability, and the traffic layer as separate design decisions with failure tests.
Database rollback restores the pre-cutover target, while Flex managed backups serve service recovery. Neither automatically redirects users to the source VM. Define write ownership, traffic-switch authority, rollback deadline, retention, and recovery objectives before migration.
The following broader design illustrates possible additions, not resources created by the reference Terraform. Additional node pools, topology rules, persistent volumes, RabbitMQ, Object Storage, and Secret Manager need their own implementation, ownership, and validation. Use them only for a demonstrated workload requirement; do not infer HA from this diagram.
cp env.tfvars.example env.tfvarsservice_account_key_path = "/path/to/stackit-sa-key.json"create_project = truetarget_project_owner_email = "owner@sa.stackit.cloud"parent_container_id = "cmf-parent-container-id"ske_cluster_name = "rpltfk8s01"observability_instance_name = "cmf-rpltf-observability"dns_zone_name = "cmf-example.runs.onstackit.cloud"dns_zone_display_name = "cmf-example"observability_enabled = truecreate_observability_instance = truedns_enabled = truecreate_dns_zone = truedeploy_workload = trueenable_postgres_flex = trueenable_springboot_hpa = falseenable_load_generator = falsespringboot_replicas = 1deploy_postgres_migration_job = falsecreate_grafana_dashboard = trueflags.env):setup_project=truesetup_observability=truesetup_database=truesetup_workload=truesetup_loadgen=falsesetup_dns=trueterraform initterraform validateterraform plan -var-file=env.tfvars -out=tfplanterraform apply tfplanExpected result: springboot_url reaches the application through the Gateway, the application
uses PostgreSQL Flex, and grafana_dashboard_url opens the managed dashboard. Provisioning
does not import source data. Follow the separate rehearsal and cutover workflow after target
validation; keep HPA disabled throughout migration.
Design the database move independently of runtime provisioning. Define source freeze, a consistent dump and manifest, isolated rehearsal, a proven target backup, transactional restore, and the rollback deadline. Traffic switching and source failback remain explicit operator decisions.
Replatform keeps core application behavior but changes selected platform components to gain operational or economic benefits. It sits between Rehost and Refactor in change intensity.
These are Replatform changes as long as the core product behavior and major code paths remain mostly stable.
Platform component selection
Identify which layers should change (for example runtime, database operations, integration controls).
Compatibility boundaries
Validate technical constraints and fallback options before introducing platform changes.
Risk-managed sequencing
Stage changes to avoid coupling too many unknowns in one cutover window.
Evidence and acceptance
Define measurable improvements for performance, resilience, and operational load.
For stateful workloads, define source and target data platform responsibilities before runtime cutover.
For a runnable example of a platform swap from VM to Kubernetes with Spring Boot, use:
Use the asset for the runnable VM-to-Kubernetes and VM-to-managed-database implementation details.
Define landing zone controls and guardrails as the start condition for the Replatform path. Confirm platform prerequisites for runtime, data, and integration layers so substitutions can be introduced without breaking governance or operability.
Define the target platform mapping for the Replatform path across runtime, data, and integration services. Make dependencies explicit, including identity, networking, and data responsibilities, so each change can be validated before cutover.
Specify required platform prerequisite changes and sequencing for controlled transition. Define rollback guardrails, readiness checks, and run ownership so wave delivery stays predictable when multiple platform layers change together.
Establish governance, identity, security, network, cost controls, and automation before the target depends on them. Confirm project permissions, operator access, DNS delegation, SKE capacity, database access boundaries, and protected Terraform state.
A landing zone is the structured cloud foundation that defines how your organization operates on STACKIT from day 1. It combines governance, identity, security, network design, cost controls, and automation into one coherent baseline.
Without this foundation, migration waves typically stall due to missing approvals, inconsistent controls, and repeated platform decisions.
The following visual summarizes the core components that should be addressed for a reliable platform baseline.
Start the landing-zone stream as early as possible, in parallel with discovery.
The practical model is a dual track: establish the platform baseline early, then refine application landing zone templates as discovery insights mature.
Platform Landing Zone
Company-wide foundation for governance, identity, security, networking, cost controls, and automation.
Open Platform Landing ZoneApplication Landing Zone
Workload-specific implementation patterns derived from the platform baseline and discovery findings.
Open Application Landing ZoneTo design a landing zone effectively, enterprises usually provide:
To accelerate delivery, STACKIT provides concrete best practices and reusable templates:


Use the Landing Zone Accelerator where appropriate for the governed project and platform foundation. Keep responsibility separate: the application team owns the SKE workload, data migration, Gateway, telemetry, and workload recovery.
Enter the migration wave with an approved target design, a ready Application Landing Zone, a tested runbook, and assigned decision owners. Follow readiness, migration, cutover, validation, stabilization, and handover as one controlled flow.
This module executes the actual migration delivery in the Migration Factory. By this point, design decisions are approved, landing zones are ready, and wave plans are defined.
Migrate focuses on repeatable technical run patterns for:
Migrate starts after Design and Mobilize has produced implementable inputs:
Reference modules:
Runbook evidence
Completed runbook records, decision logs, and rollback checkpoints for each migrated workload.
Cutover report
Time-stamped cutover outcome with acceptance results, defects, and mitigation actions.
Operational baseline
Initial monitoring, alerting, ownership, and incident procedures in the target setup.
Optimize backlog
Structured list of rightsizing, performance, and cost measures for post-cutover tuning.
Post-cutover care can overlap with Optimize after cutover, but it is handled in the Run phase.
STACKIT
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.
This runbook supplies the approval and operational sequence around the reference repository’s
scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement
are separate operations. A successful Terraform apply is not migration acceptance.
The reference migrates the public schema and validates public.album using row count and a
deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source
export. A real workload needs its own compatible export, schema assessment, business tests,
and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration,
not replication or zero-downtime cutover.
The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.
| Role | Accountable decision or evidence |
|---|---|
| Migration lead | Window, checkpoints, go/no-go authority, rollback deadline, and incident coordination |
| Application owner | All writers identified, source freeze, business acceptance, and post-cutover write reconciliation |
| Platform engineer | Approved Terraform plan, SKE access, Gateway and DNS readiness, suspended reconcilers |
| Database owner | Consistent source evidence, rehearsal, protected backup, restore integrity, and rollback execution |
| Operations owner | Telemetry, incident routing, recovery ownership, retention, and stabilization exit |
These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.
Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.
Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.
With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.
Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.python3 scripts/migrate_postgres.py rehearse \ --artifacts ../stackit-cmf-Rehost-springboot/artifacts \ --evidence .tmp/migration-runpython3 scripts/migrate_postgres.py cutover \ --artifacts ../stackit-cmf-Rehost-springboot/artifacts \ --evidence .tmp/migration-run \ --source-write-frozen --confirm-target springmusicAccept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.
| Gate | Required evidence |
|---|---|
| Data integrity | Source manifest checksum, expected count and fingerprint match the restored target; migration journal identifies the correct project and database |
| Runtime | Original replica count restored, rollout healthy, no unexplained restart or connection failures |
| Traffic | Gateway accepted, HTTPRoute references resolved, DNS matches the Gateway, and actual client requests reach the intended target |
| Business behavior | Migrated albums visible and approved user journeys pass; any write test has an agreed cleanup and reconciliation plan |
| Database transport | Application and migration connections require TLS; ACL admits only approved SKE egress or explicitly approved overrides |
| Observability | Both scrape jobs have actual up=1 samples, exporter reports database health, and dashboard values correspond to the workload |
| Recovery | Pre-cutover backup, checksum, original fingerprint, and journal retained privately and copied to protected durable storage |
| Configuration | Reviewed post-migration plan has no unexplained resource drift; source and target operating states are recorded |
Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.
Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.
Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.
python3 scripts/migrate_postgres.py rollback \ --evidence .tmp/migration-run --confirm-target springmusicReturning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.
If cutover failed before a valid backup was recorded, inspect the journal and database with the
database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed
process, inspect remaining springmusic-migration-* pods and the stopped Deployment before
resuming. The local migration lock does not coordinate different execution hosts.
Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.
Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.
Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.
Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.
| Checkpoint | Owner | Timestamp | Result | Evidence reference |
|---|---|---|---|---|
| Target and client path ready | Platform engineer | YYYY-MM-DD HH:MM | Pass/Fail | Protected record |
| Source freeze and final export | Application and DB owners | YYYY-MM-DD HH:MM | Pass/Fail | Manifest and approval |
| Final rehearsal accepted | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Rehearsal journal |
| Pre-cutover backup proven | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Backup checksum and restore result |
| Cutover and integrity accepted | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Migration journal |
| Business and traffic accepted | Application owner | YYYY-MM-DD HH:MM | Pass/Fail | Test and routing evidence |
| Drift and telemetry reviewed | Platform engineer | YYYY-MM-DD HH:MM | Pass/Fail | Plan and metric evidence |
| Handover or rollback completed | Migration lead | YYYY-MM-DD HH:MM | Pass/Fail | Signed decision |
STACKIT
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.
This runbook supplies the approval and operational sequence around the reference repository’s
scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement
are separate operations. A successful Terraform apply is not migration acceptance.
The reference migrates the public schema and validates public.album using row count and a
deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source
export. A real workload needs its own compatible export, schema assessment, business tests,
and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration,
not replication or zero-downtime cutover.
The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.
| Role | Accountable decision or evidence |
|---|---|
| Migration lead | Window, checkpoints, go/no-go authority, rollback deadline, and incident coordination |
| Application owner | All writers identified, source freeze, business acceptance, and post-cutover write reconciliation |
| Platform engineer | Approved Terraform plan, SKE access, Gateway and DNS readiness, suspended reconcilers |
| Database owner | Consistent source evidence, rehearsal, protected backup, restore integrity, and rollback execution |
| Operations owner | Telemetry, incident routing, recovery ownership, retention, and stabilization exit |
These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.
Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.
Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.
With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.
Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.python3 scripts/migrate_postgres.py rehearse \ --artifacts ../stackit-cmf-Rehost-springboot/artifacts \ --evidence .tmp/migration-runpython3 scripts/migrate_postgres.py cutover \ --artifacts ../stackit-cmf-Rehost-springboot/artifacts \ --evidence .tmp/migration-run \ --source-write-frozen --confirm-target springmusicAccept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.
| Gate | Required evidence |
|---|---|
| Data integrity | Source manifest checksum, expected count and fingerprint match the restored target; migration journal identifies the correct project and database |
| Runtime | Original replica count restored, rollout healthy, no unexplained restart or connection failures |
| Traffic | Gateway accepted, HTTPRoute references resolved, DNS matches the Gateway, and actual client requests reach the intended target |
| Business behavior | Migrated albums visible and approved user journeys pass; any write test has an agreed cleanup and reconciliation plan |
| Database transport | Application and migration connections require TLS; ACL admits only approved SKE egress or explicitly approved overrides |
| Observability | Both scrape jobs have actual up=1 samples, exporter reports database health, and dashboard values correspond to the workload |
| Recovery | Pre-cutover backup, checksum, original fingerprint, and journal retained privately and copied to protected durable storage |
| Configuration | Reviewed post-migration plan has no unexplained resource drift; source and target operating states are recorded |
Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.
Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.
Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.
python3 scripts/migrate_postgres.py rollback \ --evidence .tmp/migration-run --confirm-target springmusicReturning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.
If cutover failed before a valid backup was recorded, inspect the journal and database with the
database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed
process, inspect remaining springmusic-migration-* pods and the stopped Deployment before
resuming. The local migration lock does not coordinate different execution hosts.
Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.
Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.
Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.
Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.
| Checkpoint | Owner | Timestamp | Result | Evidence reference |
|---|---|---|---|---|
| Target and client path ready | Platform engineer | YYYY-MM-DD HH:MM | Pass/Fail | Protected record |
| Source freeze and final export | Application and DB owners | YYYY-MM-DD HH:MM | Pass/Fail | Manifest and approval |
| Final rehearsal accepted | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Rehearsal journal |
| Pre-cutover backup proven | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Backup checksum and restore result |
| Cutover and integrity accepted | DB owner | YYYY-MM-DD HH:MM | Pass/Fail | Migration journal |
| Business and traffic accepted | Application owner | YYYY-MM-DD HH:MM | Pass/Fail | Test and routing evidence |
| Drift and telemetry reviewed | Platform engineer | YYYY-MM-DD HH:MM | Pass/Fail | Plan and metric evidence |
| Handover or rollback completed | Migration lead | YYYY-MM-DD HH:MM | Pass/Fail | Signed decision |
Return to the Migration Framework's Optimize loop: collect representative operating evidence, identify the limiting layer, implement one controlled change, and validate reliability, performance, and cost before keeping it.
Optimize starts when workloads run on STACKIT and real operating data is available. The module converts post-cutover observations into measurable improvements for performance, stability, and cost efficiency.
Optimize is not a one-time task. It is an iterative cycle that can overlap with early stabilization and post-cutover care.
Many right-sizing and tuning decisions are only reliable under real load patterns. After cutover, teams can use production telemetry to separate assumptions from actual behavior.
Optimization decisions should be based on runtime evidence, not assumptions. For practical implementation, combine workload telemetry, alerting, and controlled infrastructure changes.
For Replatform workloads on Kubernetes, optimization spans multiple layers and should be coordinated as one control loop.
Primary inputs
Cutover reports, incident trends, SLO measurements, telemetry baselines, and cost reports.
Optimization outputs
Prioritized improvement backlog, validated tuning changes, and updated runbook standards.
Governance outcome
Clear trade-off decisions between performance, resilience, and cost with documented ownership.
After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.
Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.
Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repositoryOpen grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.
| Panel group | Decision supported | Interpretation boundary |
|---|---|---|
| Cluster CPU and memory | Worker pressure and aggregate capacity | CPU query reports busy cores, not a utilization percentage; cluster totals do not identify an individual pod bottleneck |
| Running pods | Workload presence | A scrape fallback is not proof that all replicas are healthy; verify Kubernetes rollout and desired replicas |
| Application requests | Request rate and mean duration | The Boot 2 adapter does not supply latency percentiles or a complete error-rate SLO |
| PostgreSQL availability and connections | Database reachability and connection pressure | Inspect pg_up and actual scrape health independently |
| PostgreSQL transactions | Commit and rollback trends | Correlate changes with traffic and application behavior |
| PostgreSQL cache hits | Read-cache behavior | Low traffic and absent series cannot establish a capacity requirement |
| PostgreSQL temp bytes and locks | Query or contention investigation | More compute is not automatically the remedy for query or lock problems |
Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some
cluster panels include fallback values, so a rendered zero is not evidence of zero consumption.
Scope queries to the intended cluster and database when a datasource contains multiple workloads.
Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.
Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.
Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.
Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.
Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.
The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas,
postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or
Replica selection resolve a flavor from the project’s current catalog. Select an offered
combination; do not assume arbitrary values or an in-place transition are supported.
Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.
The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.
PostgreSQL Flex flavors and performance classes Open the documentation| Description | ID | CPU | RAM | max_connections | shared_buffers | work_mem | maintenance_work_mem | effective_cache_size |
|---|---|---|---|---|---|---|---|---|
| Small, Compute optimized | 2.4 | 2 | 4 GB | 95 | 950 MB | 14 MB | 380 MB | 2660 MB |
| Small, Memory optimized | 2.16 | 2 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| Medium, Compute optimized | 4.8 | 4 | 8 GB | 195 | 1950 MB | 14 MB | 780 MB | 5460 MB |
| Medium, Memory optimized | 4.32 | 4 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| Large, Processor optimized | 8.16 | 8 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| X-Large, Compute optimized | 16.32 | 16 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| X-Large, Memory optimized | 16.128 | 16 | 128 GB | 3170 | 31950 MB | 14 MB | 12780 MB | 89460 MB |
max_connections limit.This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
| Description | ID | Max. IOPS | Max. throughput (MB/s) |
|---|---|---|---|
| Performance class 2 | premium-perf2-stackit | 1000 | 100 |
| Performance class 4 | premium-perf4-stackit | 2000 | 150 |
| Performance class 6 | premium-perf6-stackit | 5000 | 200 |
| Performance class 8 | premium-perf8-stackit | 10000 | 250 |
| Performance class 10 | premium-perf10-stackit | 15000 | 300 |
| Performance class 12 | premium-perf12-stackit | 20000 | 350 |
Currently, we offer three types of instances. For each type there is a different set of flavors available.
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated
CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m
and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the
two exporter sidecars has its own resource budget.
Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar
usage before editing the Deployment. Leave room beyond the Java heap for threads and native
memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and
validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.
Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.
HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.
Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.
Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:
enable_springboot_hpa = truespringboot_hpa_min_replicas = 1springboot_hpa_max_replicas = 3springboot_hpa_target_cpu_utilization_percentage = 70The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing
replica changes and establish an explicit ownership policy before unattended HPA operation.
The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.
Review the plan, then inspect HPA behavior with the configured kubeconfig:
terraform plan -var-file=env.tfvars -out=tfplan.optimizeterraform apply tfplan.optimizekubectl get hpa,pods -n springbootkubectl describe hpa springboot -n springbootkubectl top pods -n springboot --containersSupply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.
Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate
requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum
values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.
The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.
SKE node-pool management Open the documentationThe implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.
Spring Music stores its authoritative data in Flex. There is no application PersistentVolume
to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database
capacity. Use the Flex storage controls for album data and review growth, query I/O, retention,
and recovery together. Add Kubernetes storage only for a separately designed persistence need.
For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.
Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.
The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.
Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail