Skip to content
Beta

Replatform to STACKIT: Spring Boot and PostgreSQL Overview

Last updated on

Stackit LogoStackit Logo
STACKIT

Replatform to STACKIT: Spring Boot and PostgreSQL Overview

Plan the Spring Boot platform change to SKE and PostgreSQL Flex: discovery, architecture, landing-zone readiness, migration gates, handoff, and optimization.

PLAN

Spring Boot Replatform Journey

Place the platform change within the complete Migration Framework: assess the workload, design SKE and PostgreSQL Flex, prepare the landing zone, migrate with explicit data gates, and stabilize before optimization. The application JAR stays the same; the runtime and database operating models change.

STACKIT Cloud Migration Framework Journey across the four phases Assess, Design and Mobilize, Migrate, and Run with their key modules. Assess PHASE 1 Assess Design & mobilize PHASE 2 Design & mobilize Migrate PHASE 3 Migrate Run PHASE 4 Run PLANNING Discovery Discovery Map apps and dependencies. Design Design Define target patterns. Migration plan Migration plan Sequence waves for delivery. Rapid discovery Rapid discovery Inventory workloads to create an early scope and cost baseline. TCO report TCO report Model migration economics to support investment and planning decisions. Readiness assessment Readiness assessment Assess technology and organization gaps before detailed design starts. Deepdive workshop Deepdive workshop Explore STACKIT services and platform options for target design. Briefings & workshops Briefings & workshops Align stakeholders on goals, scope, and migration expectations. Business case Business case Compare value, effort, and investment per application. Enablement ENABLEMENT Center of Excellence Trainings & Learning Paths Documentation & Reference Landing zone Landing zone Establish the secure platform base for migrated workloads. Security & compliance Security & compliance Define security and compliance controls for migration and operations. Target operating model Target operating model Define roles, processes, and ownership for target operations. Migration factory setup Migration factory setup Prepare teams, tools, and runbooks for scalable migration execution. Migrate MIGRATE Relocate Relocate Rehost Rehost Replatform Replatform Optimize Optimize Improve sizing, performance, and cost after cutover. Repurchase Repurchase Evaluate SaaS options when replacement delivers better value. Refactor Refactor Restructure strategic workloads for cloud- native scalability and agility. Operating model handover Operating model handover TOM validation and ownership transfer Hypercare Hypercare Post-cutover stabilization bridge. Operate Operate Steady-state cloud operations. Support Support Manage incidents and service requests across support responsibilities. Customer Success Customer Success Track adoption, outcomes, and value realization with stakeholders.
STACKIT Migration Framework from Assess through Design and Mobilize and Migrate to Run
PLAN

Rapid Discovery

Establish an initial workload baseline for the Spring Boot VM, PostgreSQL data, dependencies, and capacity. Qualify migration intent and identify the unknowns that detailed Discovery must resolve before a platform decision.

AssessRapid discoveryOverview In 4 trails
Rapid discovery process Input data is processed through rapid discovery tooling to produce decision-ready outputs for early migration choices. Input dataCMDB & inventory exportsAsset lists, hosts, base platform factsHypervisor & cloud reportsUsage, footprint, and utilization signalsPlatform listsKubernetes, database, and OS baselinesRapid discovery toolingIngest & normalize datasetsStandardize source records and schemaClassify & aggregate assetsConsolidate by technology and quantityApply assumptionsConfidence levels and growth factorsOutput informationConsolidated asset baselineFact base for scope and planningInitial STACKIT sizingFirst capacity assumptionsPrice indication & TCO corridorEarly cost orientation for decisions
Rapid Discovery process from source collection to an initial migration baseline

Rapid Discovery provides a fast, automated baseline of the current environment across on-premises and cloud landscapes. The focus is on quantifying the existing IT portfolio in a short time window, so teams can make early migration and commercial decisions with confidence.

At this stage, quantity and distribution matter more than deep application relationships.

Rapid Discovery builds an initial inventory of infrastructure and platform assets, including:

Compute footprint

Virtual machines and host counts.

Storage baseline

Storage capacity and storage classes.

OS landscape

Operating system families and versions.

Kubernetes baseline

Kubernetes cluster counts and baseline characteristics.

Database inventory

Database engines, sizes, and instance counts.

These metrics create the first fact-based view of migration scope.

The output of Rapid Discovery is a core input for:

  • Early price indication: Build the first STACKIT-aligned cost baseline.
  • Initial TCO view: Create a first total cost of ownership corridor.
  • Target capacity assumptions: Define initial cloud capacity and landing zone requirements.

This allows program stakeholders to align on financial direction and technical baseline before detailed planning starts.

Rapid Discovery is intentionally not a full application-level analysis. It does not include deep interviews with every application owner and does not aim to fully map all runtime dependencies.

That depth is covered in the subsequent Discovery phase, where infrastructure exports are enriched with targeted assessments and owner input to build a complete application picture.

Typical input sources include exports such as spreadsheets or similar inventory files from existing environments. This phase can be accelerated with AI-assisted tooling that extracts the required baseline metrics from uploaded datasets.

Rapid Discovery therefore acts as a prerequisite for structured cost indication and for shaping a realistic target environment strategy.

Use AI-assisted discovery assets when workload descriptions and inventory inputs need to be turned into first assessment and design artifacts for expert review.

Asset title
Framework
Asset type

The following diagram shows why Rapid Discovery is performed: raw source data is processed by tooling into a decision-ready baseline that supports early price indication and initial target sizing.

Swipe sideways to see the whole diagram
Rapid discovery process Input data is processed through rapid discovery tooling to produce decision-ready outputs for early migration choices. Input dataCMDB & inventory exportsAsset lists, hosts, base platform factsHypervisor & cloud reportsUsage, footprint, and utilization signalsPlatform listsKubernetes, database, and OS baselinesRapid discovery toolingIngest & normalize datasetsStandardize source records and schemaClassify & aggregate assetsConsolidate by technology and quantityApply assumptionsConfidence levels and growth factorsOutput informationConsolidated asset baselineFact base for scope and planningInitial STACKIT sizingFirst capacity assumptionsPrice indication & TCO corridorEarly cost orientation for decisions

A robust Rapid Discovery typically follows a clear sequence:

  1. Collect data from available sources (CMDB, hypervisor exports, cloud inventories, monitoring, storage reports, database lists).
  2. Standardize and consolidate records into a unified schema.
  3. Classify assets by workload type and technical characteristics.
  4. Aggregate results for management-level decision making.
  5. Validate initial assumptions with responsible stakeholders.

The objective is not a perfect target architecture. The objective is a reliable starting point with enough accuracy for early decisions.

Result quality depends heavily on source quality. Typical issues include duplicates, outdated entries, inconsistent naming, and missing performance data.

Recommended practice for this phase:

  • Document assumptions: Keep growth rates, consolidation factors, and capacity buffers explicit.
  • Flag unclear records: Mark uncertain entries instead of removing them too early.
  • Assign confidence levels: Label findings as high, medium, or low confidence.

This keeps cost indications traceable and allows focused refinement in the subsequent Discovery phase.

Rapid Discovery provides the volume baseline for early cost modeling. Captured assets are translated into STACKIT-relevant consumption dimensions, for example:

  • Compute sizing: Use vCPU and RAM as the baseline dimensions.
  • Storage class selection: Use storage capacity and I/O characteristics.
  • Managed service options: Use database engine and size classes.
  • Platform cost estimation: Use cluster and node counts.

Combined with operating assumptions (runtime profile, availability targets, growth trajectory), this produces a solid first price indication and an initial TCO corridor.

At the end of Rapid Discovery, the following outputs should be available at a minimum:

Consolidated asset baseline

Quantities per technology domain are consolidated in one baseline.

Meaningful segmentation

Assets are segmented by criticality, environment, and modernization potential.

Traceable assumptions

Assumptions and identified data gaps are documented transparently.

Initial cost indication

Cost ranges and primary drivers are available for early planning.

Prioritized candidates

A prioritized list for deeper Discovery activities is available.

These outputs establish the working baseline for architecture, planning, and governance in the next Assess steps.

Common Rapid Discovery risks include over-simplified categorization, incomplete source systems, or overestimating data maturity.

Proven countermeasures:

  • Combine sources: Use multiple data sources instead of relying on one source.
  • Review outliers: Check very large or very old systems systematically.
  • Align perspectives: Align finance and engineering interpretation to reduce bias.

This keeps the phase fast while preserving decision quality.

The handover point is reached when quantities, technology classes, and primary cost levers are sufficiently visible and open questions are clearly documented.

In Discovery, these open items are addressed through targeted owner interviews, deeper assessments, and dependency/compliance/operations analysis to build the full application-level picture.

STEP

Discovery

Confirm Java and PostgreSQL compatibility, state handling, scheduled writers, schema dependencies, data volume and change rate, downtime tolerance, recovery objectives, and representative demand. Validate those inputs with the application and database owners before selecting the target.

Design and mobilizeDiscoveryOverview In 4 trails
Discovery source-to-decision flow Discovery separates technical and human-driven inputs, runs technical-first and human-enriched analyses, and hands over both insight streams to follow-on modules. Discovery inputsTechnical and automated discoveryInfrastructure inventoryCMDB, VM, database, middleware, storageRuntime and utilization dataCPU, memory, I/O, network and seasonalityIntegration and flow signalsNetwork paths, APIs, identity, data movementAssessment-driven human inputSecurity and compliance contextData classes, controls, audit requirementsOwner and business inputCriticality, release windows, lifecycle plansDiscovery analysis toolingTechnical-first analysesNormalize and correlateUnify records and technical identitiesDependency mappingInfer communication and couplingUtilization and sizing analysisEstimate baseline demand corridorsPreliminary segmentationCluster by stack and environmentHuman-driven analyses (application owner input)Criticality and risk calibrationValidate business impact and constraintsWave feasibility and sequencingReconcile dependencies with release windowsAssumption and gap registerTrack open points and confidenceHandover outputsTool-derived outputsDesignTarget architecture options and sizing factsLanding zonePlatform guardrails and account structure needsMigration planWave backlog, sequencing, and cutover windowsAssessment-validated outputsSecurity and complianceControl needs, data classes, remediation pointsOperating modelRole model, ownership boundaries, process impactBusiness caseValue/risk profile and modernization priorities
Discovery analysis combining technical measurements and application-owner evidence

Discovery is one of the first and most critical modules in the Design and Mobilize phase. It refines Rapid Discovery results and adds the depth needed to make architecture and migration-wave decisions with confidence.

The primary objective is to establish a realistic, evidence-based understanding of the current IT landscape, business priorities, and organizational readiness before detailed target design and migration planning are finalized.

Complete baseline

Create a reliable application and infrastructure baseline that goes beyond pure quantities.

Dependency transparency

Identify technical and process dependencies to avoid hidden migration blockers.

Business alignment

Link technical findings with business criticality, timelines, and risk tolerance.

Planning readiness

Produce decision-ready input for target design and migration-wave planning.

Inventory

Comprehensive capture of servers, virtual machines, databases, middleware, and applications.

Dependency analysis

Mapping of communication paths and runtime dependencies between systems and applications.

Resource utilization

Analysis of actual CPU, memory, storage, and I/O behavior over a representative period.

Operational context

Collection of backup, patching, SLA, compliance, and operational constraints.

Application owner input

Structured questionnaires and interviews to validate assumptions and close data gaps.

In practice, Discovery is often run together with STACKIT partners. Partners typically use their own tooling landscape to collect and normalize technical data into a central repository. Many programs also trigger targeted questionnaires for application owners directly from these tools to enrich technical findings with business and operational context.

This combined model improves speed and consistency while keeping stakeholder validation built into the process.

Two Evidence Streams: Technical vs. Human-Driven

Section titled “Two Evidence Streams: Technical vs. Human-Driven”

Discovery intentionally combines two evidence streams that complement each other:

  • Technically derived evidence: Tooling-generated findings from inventory exports, runtime metrics, and dependency signals. This stream provides scale, consistency, and repeatability.
  • Application-owner enrichment (human-driven): Validated business criticality, lifecycle intent, release constraints, and operational realities from owner interviews and questionnaires.

Neither stream is sufficient on its own. Technical evidence without owner context can misclassify critical workloads, while human input without technical grounding can hide coupling and capacity risks. Discovery quality depends on reconciling both streams into one decision-ready view.

The following diagram shows how Discovery transforms technical and stakeholder input into decision-ready outputs for the downstream modules.

Swipe sideways to see the whole diagram
Discovery source-to-decision flow Discovery separates technical and human-driven inputs, runs technical-first and human-enriched analyses, and hands over both insight streams to follow-on modules. Discovery inputsTechnical and automated discoveryInfrastructure inventoryCMDB, VM, database, middleware, storageRuntime and utilization dataCPU, memory, I/O, network and seasonalityIntegration and flow signalsNetwork paths, APIs, identity, data movementAssessment-driven human inputSecurity and compliance contextData classes, controls, audit requirementsOwner and business inputCriticality, release windows, lifecycle plansDiscovery analysis toolingTechnical-first analysesNormalize and correlateUnify records and technical identitiesDependency mappingInfer communication and couplingUtilization and sizing analysisEstimate baseline demand corridorsPreliminary segmentationCluster by stack and environmentHuman-driven analyses (application owner input)Criticality and risk calibrationValidate business impact and constraintsWave feasibility and sequencingReconcile dependencies with release windowsAssumption and gap registerTrack open points and confidenceHandover outputsTool-derived outputsDesignTarget architecture options and sizing factsLanding zonePlatform guardrails and account structure needsMigration planWave backlog, sequencing, and cutover windowsAssessment-validated outputsSecurity and complianceControl needs, data classes, remediation pointsOperating modelRole model, ownership boundaries, process impactBusiness caseValue/risk profile and modernization priorities

Typical Analysis Patterns in Discovery Tooling

Section titled “Typical Analysis Patterns in Discovery Tooling”

During Discovery, tooling commonly applies the following analysis patterns:

  • Record normalization: Merge heterogeneous exports into one coherent application model.
  • Dependency mapping: Detect communication paths, data exchange, and coupling patterns.
  • Criticality and risk scoring: Evaluate business impact, failure domain, and compliance exposure.
  • Utilization profiling: Build workload demand baselines for right-sizing and target planning.
  • Segmentation analysis: Cluster applications by readiness, constraints, and migration strategy fit.
  • Wave simulation: Model move groups and sequence options under dependency constraints.
  • Gap and assumption tracking: Keep unresolved findings transparent with confidence levels.

These analyses establish the technical fact base. The human-driven stream then validates, prioritizes, and contextualizes these findings for executable migration decisions.

Use AI-assisted discovery assets to structure workload inputs, service mapping, readiness findings, and R-strategy signals before architects validate the resulting discovery baseline.

Asset title
Framework
Asset type

  1. Aggregate source data from CMDBs, hypervisors, cloud inventories, monitoring, and export files.
  2. Normalize and consolidate records into a common application-centric model.
  3. Discover and validate dependencies (network, data, identity, integration, and batch flows).
  4. Enrich with owner input on criticality, lifecycle, constraints, and migration feasibility.
  5. Classify workloads for migration strategy options and wave sequencing.
  6. Validate findings with architecture, security, platform, and business stakeholders.
  • Reduces migration risk: Early visibility of hidden dependencies lowers outage and rollback risk.
  • Improves wave planning: Workloads can be grouped realistically by coupling, criticality, and readiness.
  • Prevents over/under-sizing: Measured utilization replaces assumptions in target capacity planning.
  • Supports governance: Security, compliance, and operational constraints are addressed before rollout.
  • Strengthens stakeholder buy-in: Shared facts improve decision quality across business and IT.

Discovery outputs are directly reused by the next modules in Design and Mobilize:

Design

Uses dependency, capacity, and risk insights to shape target architecture options.

Security and Compliance

Uses data classification and control gaps to define prioritized security requirements.

Landing Zone

Uses platform and governance constraints to define foundational setup decisions.

Migration Plan

Uses move groups, criticality, and sequencing constraints for realistic wave planning.

Operating Model and Business Case

Uses ownership, process impact, and value/risk signals for staffing and investment priorities.

At minimum, Discovery should produce the following outputs:

  • Consolidated application baseline: Mapped inventory by domain, environment, and criticality.
  • Dependency map: Verified upstream/downstream relationships and integration touchpoints.
  • Utilization profile: Evidence-based resource behavior and sizing assumptions.
  • Constraint register: Security, compliance, licensing, and operational constraints.
  • Migration readiness view: Prioritized candidates, risks, and sequencing recommendations.

These outputs are essential prerequisites for continuing with detailed design work and a credible migration plan.

LIFT

Replatform Strategy and Tools

Confirm the two deliberate substitutions: a VM service becomes a Kubernetes Deployment, and VM-local PostgreSQL becomes PostgreSQL Flex. Preserve the Spring Music JAR and business behavior; this is Replatform rather than VM Rehost or application Refactor.

Design and mobilizeDesignReplatform In 2 trails
R-strategy migration method Decision flow from discovery to production with the seven R-strategies: Relocate, Rehost, Replatform, Repurchase, Refactor, Retain, and Retire. R-strategy migration methodFrom discovery and path selection through the seven R-strategies to validation, transition, and production.DiscoveryDiscoveryAssess / prioritizeAssess / prioritizeDetermine migration pathDetermine migration pathValidationValidationTransitionTransitionProductionProductionRelocateRelocate(move VM)Define Landing ZoneDefine Landing ZoneUse migration toolsUse migration toolsAUTOMATEMANUALInstallInstallConfigConfigDeployDeployValidation & handoverRehostingRehosting(move application)Define Landing ZoneDefine Landing ZoneUse migration toolsUse migration toolsAUTOMATEMANUALInstallInstallConfigConfigDeployDeployReplatformingReplatforming(lift and reshape)Define Landing ZoneDefine Landing ZoneMap Target PlatformMap Target PlatformAdapt Platform StackAdapt Platform StackRepurchasingRepurchasing(replace, drop and shop)Purchase COTS/SaaS and licensingPurchase COTS/SaaS and licensingMigrate business processMigrate business processRefactoringRefactoring(re-architecting applications)Redesign application/ infrastructure architectureRedesign application/ infrastructure architectureApp code developmentApp code developmentFull ALM/SDLCFull ALM/SDLCIntegrationIntegrationRetain/moveRetain/movekeep for now or move laterRetire/decommissionRetire/decommissionLanding zone foundationLanding zone foundationShared platform base for all paths
R-strategy method placing Replatform between Rehost and Refactor

Replatform keeps core application behavior but changes selected platform components to gain operational or economic benefits. It sits between Rehost and Refactor in change intensity.

Comparing Replatform, Rehost, and Refactor

Section titled “Comparing Replatform, Rehost, and Refactor”
  • Rehost: Move workload location with minimal platform or code change. Example: Spring Boot stays on VM, only cloud target changes.
  • Replatform: Keep application behavior, but change selected platform layers. Example: Spring Boot runtime moves from VM to Kubernetes while core business logic remains unchanged.
  • Refactor: Change code structure or architecture significantly to unlock additional capabilities. Example: split monolith into services, redesign persistence model, and rework integration contracts.
  • Runtime platform swap: VM-based application hosting to Kubernetes.
  • Data platform swap: Self-managed database on VM to managed PaaS database service.
  • Ops capability swap: Host-centric monitoring and deployment model to managed platform-native operations.
  • Connectivity/control swap: Ingress, DNS, and service exposure model adapted to managed platform patterns.

These are Replatform changes as long as the core product behavior and major code paths remain mostly stable.

  • Operational bottlenecks can be reduced through managed platform capabilities.
  • Moderate change tolerance exists, but full redesign is out of scope.
  • Scalability and reliability goals require infrastructure-level improvements.
  • Cost optimization target can be reached with selective platform substitution.

Platform component selection

Identify which layers should change (for example runtime, database operations, integration controls).

Compatibility boundaries

Validate technical constraints and fallback options before introducing platform changes.

Risk-managed sequencing

Stage changes to avoid coupling too many unknowns in one cutover window.

Evidence and acceptance

Define measurable improvements for performance, resilience, and operational load.

  1. Define Landing Zone for the workload and its control boundaries.
  2. Map Target Platform components for runtime, data, and integration.
  3. Adapt Platform Stack prerequisites with sequencing and rollback checkpoints.
  4. Use migration tools to execute the transition through the shared migration path.
  5. Validate non-functional requirements and approve handover.

For stateful workloads, define source and target data platform responsibilities before runtime cutover.

  • Source data ownership: Clarify who owns dump/export run and consistency checks.
  • Target data ownership: Clarify who owns managed database provisioning, access controls, and backup baseline.
  • Migration sequencing: Separate schema/data move from runtime switch and validate each gate independently.
  • Temporary access controls: Plan temporary data migration access and explicit rollback/removal checkpoints.
  • Approved design decision record with scope, assumptions, and governance sign-off.
  • Validation evidence package for security, compliance, and operational readiness.
  • Strategy-specific migration runbook draft from the Design phase.
  • Handover package for Migration Factory Setup and wave planning.
  • Replatform decision matrix with selected substitutions.
  • Compatibility and constraint assessment.
  • Sequenced migration and rollback design.
  • Target-state operations model.
  • Benefit metrics and acceptance criteria.

For a runnable example of a platform swap from VM to Kubernetes with Spring Boot, use:

Asset title
Framework
Asset type

Use the asset for the runnable VM-to-Kubernetes and VM-to-managed-database implementation details.

Define Landing Zone

Define landing zone controls and guardrails as the start condition for the Replatform path. Confirm platform prerequisites for runtime, data, and integration layers so substitutions can be introduced without breaking governance or operability.

Map Target Platform

Define the target platform mapping for the Replatform path across runtime, data, and integration services. Make dependencies explicit, including identity, networking, and data responsibilities, so each change can be validated before cutover.

Adapt Platform Stack

Specify required platform prerequisite changes and sequencing for controlled transition. Define rollback guardrails, readiness checks, and run ownership so wave delivery stays predictable when multiple platform layers change together.

STEP

Automate Platform and Workload

Use versioned Terraform for infrastructure and Kubernetes resources, Helm for Envoy Gateway and routes, and a separate approved script for data migration. Keep provisioning and data replacement independently reviewable; this target does not require Ansible host configuration.

Design and mobilizeLanding zonesAutomation (IaC) In 3 trails

Automation ensures that landing-zone capabilities are reproducible, versioned, and tested instead of manually configured.

For migration landing zones, automation is the delivery backbone that connects platform APIs, IaC tools, developer workflows, and release controls into one reliable operating model.

  • STACKIT API: Use the API as the foundational control surface for platform automation and integration patterns. Documentation
  • Terraform Provider: Use the official provider for declarative infrastructure provisioning and lifecycle control. Documentation
  • OpenTofu Provider: Use OpenTofu with the STACKIT provider as an open IaC option with comparable declarative workflows. Documentation
  • Pulumi: Use Pulumi when teams prefer general-purpose languages for infrastructure automation. Documentation
  • Ansible: Use Ansible primarily for post-provisioning configuration and operational tasks. In a combined model, Terraform/OpenTofu provision infrastructure while Ansible applies OS and middleware configuration.
  • STACKIT CLI: Standardize CLI-based operations for scripting, troubleshooting, and repeatable operational run tasks. Documentation
  • SDKs (Go, Python, Java): Use SDKs for custom automation and service integrations where IaC abstractions are not sufficient. Go SDK , Python SDK , Java SDK .
  • STACKIT Git: Use Git as the source of truth for IaC modules, policies, and delivery workflows. Documentation
  • CI/CD Pipeline: Use pipelines for validation, policy checks, controlled promotion, and auditable releases. Documentation
  • Container Registry: Use a central registry for versioned build artifacts and deployment consistency across environments. Documentation
  • Control layer: STACKIT API, CLI, and SDKs provide direct and programmable control interfaces.
  • Provisioning layer: Terraform/OpenTofu and Pulumi define and reconcile desired infrastructure state.
  • Configuration layer: Ansible applies host and middleware configuration after infrastructure provisioning.
  • Delivery layer: Git and CI/CD pipelines enforce quality gates, policy checks, and controlled rollout across environments.
  • Artifact layer: Container Registry delivers immutable and versioned artifacts for predictable deployments.

Terraform/OpenTofu and Ansible delivery flow

Section titled “Terraform/OpenTofu and Ansible delivery flow”

Terraform or OpenTofu and Ansible solve different parts of one delivery workflow. Keep the boundary explicit so infrastructure changes remain reviewable and host configuration remains repeatable.

Terraform / OpenTofu

Own the infrastructure lifecycle: projects, networks, security controls, compute, storage, managed services, and the outputs required by configuration management.

Ansible

Own configuration inside the reachable target: operating-system packages, middleware, application artifacts, service units, and workload-level validation.

  1. Version infrastructure inputs, configuration, and application artifact references in Git.
  2. Validate and review the Terraform/OpenTofu plan, including replacement and security effects.
  3. Apply the approved plan and expose only the target inventory and outputs required by Ansible.
  4. Run Ansible idempotently to configure the operating system, middleware, workload, and telemetry.
  5. Validate infrastructure state, service health, and operational controls, then retain the evidence.
  6. Promote the same versioned workflow through environments instead of repeating manual setup.

Do not use provisioners or ad hoc scripts to blur ownership between both layers. Triggering Ansible from Terraform can be a practical bridge, but each tool must remain independently understandable, testable, and rerunnable.

  • Recommendation 1: Use Git plus CI/CD as default control path and avoid direct manual changes in productive scopes.
  • Recommendation 2: Choose one primary IaC engine per platform domain (Terraform or OpenTofu) to reduce fragmentation.
  • Recommendation 3: Use Ansible for configuration management, not as a replacement for declarative infrastructure provisioning.
  • Recommendation 4: Use SDKs for domain-specific automation where provider resources do not cover required behavior.
  • Recommendation 5: Version and promote container artifacts through clear environment stages with rollback-ready tags.
  • Automation interface strategy: Define where API, CLI, SDK, and IaC tools are used as primary interfaces.
  • IaC engine strategy: Decide Terraform versus OpenTofu versus Pulumi based on skills, governance, and ecosystem fit.
  • Module and repository strategy: Define reusable module boundaries, versioning, and ownership.
  • Pipeline control model: Implement validation, policy checks, approvals, and promotion gates.
  • Artifact and release strategy: Define registry usage, image versioning, and rollback standards.
  • Reusable automation baseline: IaC modules, templates, and configuration playbooks with ownership model.
  • Delivery blueprint: CI/CD flow with quality gates, policy checks, and staged promotions.
  • Integration toolkit: Standardized use of CLI and SDK automation for operational and product-specific workflows.
  • Release baseline: Versioned artifact lifecycle in Container Registry with rollback-ready practices.
  • Too many automation paradigms: Parallel tool stacks with no clear ownership or governance model.
  • Provisioning and configuration mixed ad hoc: No clean boundary between IaC provisioning and Ansible configuration.
  • No artifact discipline: Mutable container tags and unclear release traceability.
  • CLI scripts without Git and pipeline controls: Operational automation cannot be audited or reproduced reliably.
OPS

Target Runtime Architecture

This architecture maps the VM-based Spring Boot and PostgreSQL source to a Kubernetes runtime and managed database on STACKIT. The same application JAR is retained while provisioning, deployment, traffic management, data recovery, and operational responsibilities change.

The reference baseline uses one SKE worker and PostgreSQL Flex, with Envoy Gateway, STACKIT DNS, and Observability. It does not deploy the additional services or multi-zone topology shown in the optional extension pattern below.

  • Runtime standardization: replace a systemd-managed Java process with a reproducible Deployment and health checks.
  • Database operations: move PostgreSQL to a managed service without redesigning the application schema.
  • Controlled platform change: qualify rollout, scaling, network access, and recovery independently before production acceptance.
Source VMApproved dump + manifestApplication clientsSTACKIT Application ProjectSpring Music JAR + systemdSelf-managed PostgreSQLSTACKIT DNSSKE: single-worker referencePostgreSQL FlexObservability + GrafanaEnvoy Gateway + HTTPRoutesClusterIP ServiceSame JAR on Java 11Boot 2 adapter + PG exporterTemporary migration clientManaged ExternalDNSspringmusicspringmusic_rehearsal local SQL watch route hostnamesJDBC / TLSrehearse / prove backupapproved cutover / rollbackdatabase metrics / TLSpublish Gateway addressscrape via Gateway 9090 / 9187freeze / export / verifyprotected transfer via kubectlresolve hostnameHTTP baseline; HTTPS optional

An init container verifies the commit-pinned JAR checksum before Java starts. Application containers are replaceable: authoritative album data lives in PostgreSQL Flex, not in a pod filesystem or Kubernetes PersistentVolume. Kubernetes Secrets inject database credentials; an external Secret Manager integration is not implemented in this baseline.

The Flex ACL defaults to actual SKE egress CIDRs. Both application and migration client require encrypted database connections. The migration client uses an isolated rehearsal database and only replaces the application data after explicit approval and a verified pre-cutover backup. No source-VM database connection or temporary public Flex ACL is required for the dump-based path.

Terraform installs Envoy Gateway and then a local routing chart. The application Service is ClusterIP; Envoy supplies the public LoadBalancer. SKE-managed ExternalDNS publishes the HTTPRoute hostname from the Gateway address. This is Gateway API, not a legacy Ingress controller or a separately provisioned STACKIT Application Load Balancer service.

HTTP is the tested default. For HTTPS, supply a trusted TLS Secret and configure gateway_tls_secret_name according to the repository procedure; certificate issuance and renewal remain external responsibilities. The separate metrics listeners are public and unauthenticated in the reference and require protection before sensitive use.

Boot 2 Actuator binds to pod-local loopback; the metrics adapter exposes selected measurements. The PostgreSQL exporter and the SKE monitoring integration feed Observability. Terraform creates the Grafana folder and dashboard, but dashboard availability alone does not establish application health, scrape continuity, or working alert delivery.

The tested worker count, HTTP endpoint, and sample application are a functional baseline, not an HA production architecture. Select a supported SKE release and suitable zone capacity. Assess multiple workers, zone distribution, workload disruption budgets, replica safety, database availability, and the traffic layer as separate design decisions with failure tests.

Database rollback restores the pre-cutover target, while Flex managed backups serve service recovery. Neither automatically redirects users to the source VM. Define write ownership, traffic-switch authority, rollback deadline, retention, and recovery objectives before migration.

The following broader design illustrates possible additions, not resources created by the reference Terraform. Additional node pools, topology rules, persistent volumes, RabbitMQ, Object Storage, and Secret Manager need their own implementation, ownership, and validation. Use them only for a demonstrated workload requirement; do not infer HA from this diagram.

InternetApplication ProjectBackend ServicesAccessKubernetes (SKE)PostgreSQLRabbitMQObject StorageSecret ManagerObservabilityExternal LBDNSEntry LayerService LayerWorkload LayerPlatform LayerGateway APIExternalDNSK8s ServicePersistent StorageDeploymentHPANode Pool AZ-1Node Pool AZ-2Node AutoscalerPVPVPod APod BVMVM
  • Decouple runtime and data migration gates: validate database migration and runtime rollout independently.
  • Design secret delivery explicitly: the reference uses Kubernetes Secrets and protected Terraform state; add a reviewed external secret integration when required.
  • Standardize observability labels and dashboards: make cross-application operation and incident handling consistent.
  • Keep RabbitMQ optional and explicit: add it when asynchronous integration or buffering is required.
  • Qualify availability separately: a multi-zone design needs suitable worker capacity, placement rules, disruption budgets, and application and database failure testing; it is not enabled by the baseline.
Cloud Framework Replatform Spring Boot with Terraform Follow the executable provisioning, source-evidence, rehearsal, cutover, and rollback workflow for this architecture. Open page Code & registry github.com STACKIT CMF Replatform Spring Boot Kubernetes repository Open the repository
  1. Copy the example file: cp env.tfvars.example env.tfvars
  2. Set required identity/project values:
service_account_key_path = "/path/to/stackit-sa-key.json"
create_project = true
target_project_owner_email = "owner@sa.stackit.cloud"
parent_container_id = "cmf-parent-container-id"
ske_cluster_name = "rpltfk8s01"
observability_instance_name = "cmf-rpltf-observability"
dns_zone_name = "cmf-example.runs.onstackit.cloud"
dns_zone_display_name = "cmf-example"
  1. Enable the target architecture switches:
observability_enabled = true
create_observability_instance = true
dns_enabled = true
create_dns_zone = true
deploy_workload = true
enable_postgres_flex = true
enable_springboot_hpa = false
enable_load_generator = false
springboot_replicas = 1
deploy_postgres_migration_job = false
create_grafana_dashboard = true
  1. Optional CMF flag wrapper (flags.env):
setup_project=true
setup_observability=true
setup_database=true
setup_workload=true
setup_loadgen=false
setup_dns=true
  1. Apply:
Terminal window
terraform init
terraform validate
terraform plan -var-file=env.tfvars -out=tfplan
terraform apply tfplan

Expected result: springboot_url reaches the application through the Gateway, the application uses PostgreSQL Flex, and grafana_dashboard_url opens the managed dashboard. Provisioning does not import source data. Follow the separate rehearsal and cutover workflow after target validation; keep HPA disabled throughout migration.

Code & registry github.com Implemented topology and prerequisites Review the exact resource definitions and operational boundaries in the Spring Boot Replatform repository. Open the repository
AUTO

Database Migration

Design the database move independently of runtime provisioning. Define source freeze, a consistent dump and manifest, isolated rehearsal, a proven target backup, transactional restore, and the rollback deadline. Traffic switching and source failback remain explicit operator decisions.

Design and mobilizeDesignReplatform In 2 trails
End-to-End Migration Wave Flow Four sequential stages lead from wave approval through preparation and cutover to stabilization and handover. Every wave follows the same controlled factory flow. Secure scope and baseline, execute migration, prove acceptance, and hand over safely to operations. 1 APPROVE Scope and baseline Window, rollback, and ownership Freeze versions and dependencies 2 PREPARE Readiness and R-path Validate source and target Run playbook by archetype 3 CUT OVER Cutover and acceptance Route traffic to STACKIT safely Validate tech, function, operations 4 STABILIZE Learn and hand over Resolve findings in short loops Hand over to Optimize and Operate Traceable wave completion: accepted, stabilized, and fully handed over
Migration-wave control points from readiness through cutover, validation, and handover

Replatform keeps core application behavior but changes selected platform components to gain operational or economic benefits. It sits between Rehost and Refactor in change intensity.

Comparing Replatform, Rehost, and Refactor

Section titled “Comparing Replatform, Rehost, and Refactor”
  • Rehost: Move workload location with minimal platform or code change. Example: Spring Boot stays on VM, only cloud target changes.
  • Replatform: Keep application behavior, but change selected platform layers. Example: Spring Boot runtime moves from VM to Kubernetes while core business logic remains unchanged.
  • Refactor: Change code structure or architecture significantly to unlock additional capabilities. Example: split monolith into services, redesign persistence model, and rework integration contracts.
  • Runtime platform swap: VM-based application hosting to Kubernetes.
  • Data platform swap: Self-managed database on VM to managed PaaS database service.
  • Ops capability swap: Host-centric monitoring and deployment model to managed platform-native operations.
  • Connectivity/control swap: Ingress, DNS, and service exposure model adapted to managed platform patterns.

These are Replatform changes as long as the core product behavior and major code paths remain mostly stable.

  • Operational bottlenecks can be reduced through managed platform capabilities.
  • Moderate change tolerance exists, but full redesign is out of scope.
  • Scalability and reliability goals require infrastructure-level improvements.
  • Cost optimization target can be reached with selective platform substitution.

Platform component selection

Identify which layers should change (for example runtime, database operations, integration controls).

Compatibility boundaries

Validate technical constraints and fallback options before introducing platform changes.

Risk-managed sequencing

Stage changes to avoid coupling too many unknowns in one cutover window.

Evidence and acceptance

Define measurable improvements for performance, resilience, and operational load.

  1. Define Landing Zone for the workload and its control boundaries.
  2. Map Target Platform components for runtime, data, and integration.
  3. Adapt Platform Stack prerequisites with sequencing and rollback checkpoints.
  4. Use migration tools to execute the transition through the shared migration path.
  5. Validate non-functional requirements and approve handover.

For stateful workloads, define source and target data platform responsibilities before runtime cutover.

  • Source data ownership: Clarify who owns dump/export run and consistency checks.
  • Target data ownership: Clarify who owns managed database provisioning, access controls, and backup baseline.
  • Migration sequencing: Separate schema/data move from runtime switch and validate each gate independently.
  • Temporary access controls: Plan temporary data migration access and explicit rollback/removal checkpoints.
  • Approved design decision record with scope, assumptions, and governance sign-off.
  • Validation evidence package for security, compliance, and operational readiness.
  • Strategy-specific migration runbook draft from the Design phase.
  • Handover package for Migration Factory Setup and wave planning.
  • Replatform decision matrix with selected substitutions.
  • Compatibility and constraint assessment.
  • Sequenced migration and rollback design.
  • Target-state operations model.
  • Benefit metrics and acceptance criteria.

For a runnable example of a platform swap from VM to Kubernetes with Spring Boot, use:

Asset title
Framework
Asset type

Use the asset for the runnable VM-to-Kubernetes and VM-to-managed-database implementation details.

Define Landing Zone

Define landing zone controls and guardrails as the start condition for the Replatform path. Confirm platform prerequisites for runtime, data, and integration layers so substitutions can be introduced without breaking governance or operability.

Map Target Platform

Define the target platform mapping for the Replatform path across runtime, data, and integration services. Make dependencies explicit, including identity, networking, and data responsibilities, so each change can be validated before cutover.

Adapt Platform Stack

Specify required platform prerequisite changes and sequencing for controlled transition. Define rollback guardrails, readiness checks, and run ownership so wave delivery stays predictable when multiple platform layers change together.

BASE

Landing Zone

Establish governance, identity, security, network, cost controls, and automation before the target depends on them. Confirm project permissions, operator access, DNS delegation, SKE capacity, database access boundaries, and protected Terraform state.

Design and mobilizeLanding zonesOverview In 5 trails
Landing zone core components Six building blocks of a secure landing zone. Landing zone core componentsThe six building blocks of a secure platform foundation on STACKIT.Account GovernanceAccount GovernanceHow do I structure my projects?Identity & Access ManagementIdentity & Access ManagementWho can do what on the platform?Security & ComplianceSecurity & ComplianceHow do I monitor and protect?Network ArchitectureNetwork ArchitectureHow are components securely connected?Cost Management and ControlCost Management and ControlHow do I keep spending in control?Automation (IaC)Automation (IaC)How is everything delivered and managed?
Six core components of a governed STACKIT landing zone

A landing zone is the structured cloud foundation that defines how your organization operates on STACKIT from day 1. It combines governance, identity, security, network design, cost controls, and automation into one coherent baseline.

Without this foundation, migration waves typically stall due to missing approvals, inconsistent controls, and repeated platform decisions.

The following visual summarizes the core components that should be addressed for a reliable platform baseline.

Landing zone core components Six building blocks of a secure landing zone. Landing zone core componentsThe six building blocks of a secure platform foundation on STACKIT.Account GovernanceAccount GovernanceHow do I structure my projects?Identity & Access ManagementIdentity & Access ManagementWho can do what on the platform?Security & ComplianceSecurity & ComplianceHow do I monitor and protect?Network ArchitectureNetwork ArchitectureHow are components securely connected?Cost Management and ControlCost Management and ControlHow do I keep spending in control?Automation (IaC)Automation (IaC)How is everything delivered and managed?
  • Control and risk reduction: Enforce security and compliance controls consistently across teams.
  • Scalable delivery baseline: Enable repeatable provisioning patterns for multiple migration waves.
  • Clear responsibilities: Define ownership boundaries for platform, security, and application teams.
  • Faster migration throughput: Avoid redesigning core controls for each application move.

Start the landing-zone stream as early as possible, in parallel with discovery.

  • Too late: Productive migrations are blocked because mandatory controls are not yet available.
  • Too early without discovery feedback: Application constraints are missed and later cause rework.

The practical model is a dual track: establish the platform baseline early, then refine application landing zone templates as discovery insights mature.

Two layers: Platform and Application Landing Zones

Section titled “Two layers: Platform and Application Landing Zones”

Platform Landing Zone

Company-wide foundation for governance, identity, security, networking, cost controls, and automation.

Open Platform Landing Zone

Application Landing Zone

Workload-specific implementation patterns derived from the platform baseline and discovery findings.

Open Application Landing Zone

To design a landing zone effectively, enterprises usually provide:

  • Organization and ownership model: Entities, project boundaries, and accountability model.
  • Compliance and policy requirements: Regulatory obligations and internal control policies.
  • Security requirements: IAM standards, network segmentation, encryption, and logging expectations.
  • Operations and support constraints: Incident handling, escalation paths, and handover model.
  • Application portfolio insights: Discovery findings about workload archetypes and dependencies.
  1. Define enterprise guardrails and target control model.
  2. Build and validate the platform landing zone baseline as code.
  3. Derive application landing zone templates from discovery and migration design.
  4. Pilot with selected workloads, then scale through migration factory runbooks.

To accelerate delivery, STACKIT provides concrete best practices and reusable templates:

Asset title
Framework
Asset type

Accelerate the Foundation as Code

Use the Landing Zone Accelerator where appropriate for the governed project and platform foundation. Keep responsibility separate: the application team owns the SKE workload, data migration, Gateway, telemetry, and workload recovery.

Landing Zone Accelerator separating platform foundation and application workload responsibilities
Landing Zone Accelerator separating platform foundation and application workload responsibilities
LIVE

Migration Framework

Enter the migration wave with an approved target design, a ready Application Landing Zone, a tested runbook, and assigned decision owners. Follow readiness, migration, cutover, validation, stabilization, and handover as one controlled flow.

MigrateMigrateOverview In 4 trails

This module executes the actual migration delivery in the Migration Factory. By this point, design decisions are approved, landing zones are ready, and wave plans are defined.

Migrate focuses on repeatable technical run patterns for:

  • Relocate: Move workloads with minimal change in platform assumptions.
  • Rehost: Lift and shift workloads to STACKIT with controlled cutover and rollback readiness.
  • Replatform: Apply targeted platform changes during migration without a full application redesign.

Migrate starts after Design and Mobilize has produced implementable inputs:

  • Application design per workload: Target-state scope, interface decisions, data handling, and non-functional requirements are defined.
  • Wave assignment: Every workload is assigned to an approved migration wave with timing and dependency logic.
  • Runbook availability: A validated runbook exists per migration path and workload archetype, including go/no-go, cutover, and rollback criteria.
  • Application Landing Zone mapping: Workloads are mapped to the correct target landing zone and ownership model.
  • Decision and communication model: Decision paths, escalation process, and communication with the Application Owner are defined, including whether cutover must be performed jointly.

Reference modules:

  1. Confirm wave scope, cutover window, rollback criteria, and runbook ownership.
  2. Freeze wave baseline (application version, dependencies, data scope, and interface contracts).
  3. Run pre-migration technical readiness checks in source and target environments.
  4. Run migration runbook steps per workload archetype and R-strategy path.
  5. Perform cutover and route traffic to the STACKIT target according to the wave plan.
  6. Validate technical, functional, and operational acceptance criteria.
  7. Stabilize incidents and known defects in short feedback loops.
  8. Hand over stabilized workloads to Optimize and Operate.
  1. Recreate workload placement in STACKIT with equivalent topology assumptions.
  2. Migrate data and state with consistency checks.
  3. Cut over traffic with rollback guardrails.
  4. Validate business continuity and operational telemetry.
  1. Move application and data components largely unchanged.
  2. Reconfigure infrastructure bindings and connectivity on STACKIT.
  3. Run controlled cutover and smoke tests.
  4. Complete stabilization and baseline performance validation.
  1. Introduce selected managed services or platform capabilities during migration.
  2. Adapt configuration, deployment packaging, and operational controls.
  3. Cut over with compatibility and data integrity checks.
  4. Validate SLO behavior and operating model readiness.

Runbook evidence

Completed runbook records, decision logs, and rollback checkpoints for each migrated workload.

Cutover report

Time-stamped cutover outcome with acceptance results, defects, and mitigation actions.

Operational baseline

Initial monitoring, alerting, ownership, and incident procedures in the target setup.

Optimize backlog

Structured list of rightsizing, performance, and cost measures for post-cutover tuning.

Boundary to optimize, repurchase, and refactor

Section titled “Boundary to optimize, repurchase, and refactor”
  • Optimize starts directly after stable cutover and uses production telemetry to tune performance and cost: Optimize.
  • Repurchase follows a different SaaS transition logic and is treated as a dedicated module outside wave-based factory mechanics: Repurchase.
  • Refactor is typically run as a separate transformation stream: Refactor.

Post-cutover care can overlap with Optimize after cutover, but it is handled in the Run phase.

SAFE

Migration Decisions and Gates

Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.

This runbook supplies the approval and operational sequence around the reference repository’s scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement are separate operations. A successful Terraform apply is not migration acceptance.

The reference migrates the public schema and validates public.album using row count and a deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source export. A real workload needs its own compatible export, schema assessment, business tests, and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration, not replication or zero-downtime cutover.

The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.

  • Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
  • Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
  • Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
  • Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
  • Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.

These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.

Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.

  • Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
  • Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
  • Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
  • Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
  • Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
  • Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
  • Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.

Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.

From the STACKIT docsCreate and manage instances › ACLSource updated 24.09.2026 · copied 06.10.2026

With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.

Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
  2. Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
  3. Run bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.
  4. Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
  5. Set enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
  1. Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
  2. Export the final consistent dump and manifest through the approved source procedure.
  3. Execute rehearsal from the Replatform repository using the approved artifact directory:
Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run
  1. Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
  2. Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
  1. Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
  2. Execute the explicitly confirmed cutover:
Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic
  1. Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
  2. Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
  3. Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
  4. Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.

Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.

Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.

Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.

Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.

  1. Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
  2. Restore the protected pre-cutover target using the original evidence directory:
Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic
  1. Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
  2. Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
  3. Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.

Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.

If cutover failed before a valid backup was recorded, inspect the journal and database with the database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed process, inspect remaining springmusic-migration-* pods and the stopped Deployment before resuming. The local migration lock does not coordinate different execution hosts.

Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.

Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.

Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.

Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.

Code & registry github.com Executable migration and recovery workflow Use the reference repository for exact command prerequisites, evidence formats, safeguards, and validation coverage. Open the repository
GOAL

Accept and Hand Over to Operations

Move the Spring Music application from a VM to STACKIT Kubernetes Engine and its data from self-managed PostgreSQL to PostgreSQL Flex. Preserve the application JAR and business behavior while introducing Kubernetes deployment, Gateway API, DNS, and managed observability.

This runbook supplies the approval and operational sequence around the reference repository’s scripts/migrate_postgres.py commands. Infrastructure provisioning and database replacement are separate operations. A successful Terraform apply is not migration acceptance.

The reference migrates the public schema and validates public.album using row count and a deterministic fingerprint. The tested input is the Rehost eight-album sample, not a live-source export. A real workload needs its own compatible export, schema assessment, business tests, and recovery objectives. Approve downtime: this is a write-freeze and dump/restore migration, not replication or zero-downtime cutover.

The target uses a dedicated application database and rehearsal database, TLS-required database connections, and a temporary in-cluster migration client. The script does not stop source writers, switch client traffic, configure public TLS, or automate source failback. Those are operator tasks.

  • Ready to migrate: approve the target design, access boundaries, downtime, responsibilities, and measurable acceptance criteria before opening the window.
  • Ready to cut over: freeze source writes and require a consistent final dump with matching, recent rehearsal evidence. Infrastructure readiness alone does not authorize replacing data.
  • Protected execution: exclude competing writers and reconcilers, prove the pre-cutover target backup, and validate the transactional restore before restarting the workload.
  • Accept or recover: require matching data, successful business journeys, working client traffic, and actual telemetry. Decide rollback before the deadline; source failback and post-cutover write reconciliation remain separate decisions.
  • Ready for operations: transfer evidence and recovery ownership, stabilize the workload, and resume automation deliberately. Capacity and HPA experiments belong to a later change window.

These gates explain the control model. The following sections provide the executable procedure and evidence requirements for the technical walkthrough.

Confirm writer control, paused reconcilers, ownership, acceptance criteria, and the rollback deadline.

  • Source baseline: record JAR checksum, Java and PostgreSQL versions, schema dependencies, data size and change rate, scheduled jobs, integrations, and recovery objectives.
  • Source evidence: validate the trusted dump and manifest from one consistent snapshot; record checksum, expected row count, and fingerprint. Rehearse the final dump after the write freeze.
  • Target readiness: complete the reviewed infrastructure apply, check SKE capacity, artifact access, Flex connectivity and ACLs, Gateway conditions, DNS, application responses, and both metrics jobs.
  • Access and security: verify operator kubeconfig and permissions, protect state and plans, restrict secrets and evidence, and resolve HTTP and public-metrics limitations for the intended data classification.
  • Writer control: disable HPA and load generation, stop other target writers, and suspend Terraform, GitOps, and scheduled deployment jobs during migration. Reserve the target for one operator workflow.
  • Recovery readiness: agree target identity, evidence location, protected off-container backup storage, rollback authority, deadline, traffic-switch procedure, and source retention.
  • Acceptance: define permitted downtime, data invariants, business tests, error and latency thresholds, and the response to missing telemetry before the window starts.

Before entering the window, verify the PostgreSQL Flex ACL against the actual migration-client and application source addresses. Network admission is an additional control, not a replacement for database authentication or the TLS-required connections used by this runbook.

From the STACKIT docsCreate and manage instances › ACLSource updated 24.09.2026 · copied 06.10.2026

With the ACL entries, you control which source IPs are allowed to connect to your instance. Note, that this is an additional security layer and does not replace the need for proper authentication and security best practices. There are two predefined entries: 193.148.160.0/19 and 45.129.40.0/21. They ensure that you can access your instance from STACKIT cloud services. If you want to access your instance from the public net, you need to add the client’s IPv4 address or subnet. The entries follow the CIDR notation. If you want to allow a single IP address (e.g. single host), then set 32as the subnet parameter. E.g. to allow a host with the source IPv4 address of 93.229.84.137, add 93.229.84.137/32 as ACL entry. At the moment, you can’t add IPv6 addresses.

Do not set 0.0.0.0/0 as an ACL IP, because then your instance can be accessed from every IP.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Confirm the approved code revision, variable file, project, cluster, namespace, and target database.
  2. Provision the target through a reviewed saved Terraform plan; reject unrelated resource replacements.
  3. Run bash scripts/validate_gateway.sh and inspect workload rollout and PostgreSQL connectivity.
  4. Capture source evidence and baseline target behavior. A seed-data target is not an accepted migrated target.
  5. Set enable_springboot_hpa = false, enable_load_generator = false, and deploy_postgres_migration_job = false; apply those settings before suspending infrastructure automation.
  1. Obtain go/no-go approval and freeze all source writers, including integrations and background jobs.
  2. Export the final consistent dump and manifest through the approved source procedure.
  3. Execute rehearsal from the Replatform repository using the approved artifact directory:
Terminal window
python3 scripts/migrate_postgres.py rehearse \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run
  1. Require matching checksum, row count, fingerprint, and target identity; verify the application database remains unchanged.
  2. Confirm successful rehearsal is less than 24 hours old and the final dump has not changed. Reject stale or mismatched evidence.
  1. Reconfirm source freeze, target-writer exclusion, paused reconcilers, the decision deadline, and available evidence storage.
  2. Execute the explicitly confirmed cutover:
Terminal window
python3 scripts/migrate_postgres.py cutover \
--artifacts ../stackit-cmf-Rehost-springboot/artifacts \
--evidence .tmp/migration-run \
--source-write-frozen --confirm-target springmusic
  1. Require the script to stop application replicas, save the pre-cutover target dump, and prove its restore in the rehearsal database before importing the source.
  2. Require the transactional source restore and data checks to succeed before the script restores the original replica count. Investigate a stopped application after any failure; do not override it with Terraform.
  3. Complete technical and business validation, then switch client traffic through the approved operator procedure. Verify both the new client path and the continued source write freeze.
  4. Record acceptance and write ownership. Once the script has completed and replicas are correct, review a Terraform plan for drift and explain any output-only kubeconfig refresh; do not apply unrelated changes during acceptance.

Accept only with matching data evidence, working client traffic, healthy runtime, and actual telemetry.

Public HTTP success alone does not satisfy a production HTTPS requirement. The sample dashboard does not replace independent business, latency-percentile, error-rate, or recovery validation.

Restore the protected target when an approved trigger is met; reconcile post-cutover writes and decide source failback separately.

Invoke the agreed decision before the deadline when data invariants fail, a critical business journey cannot be restored within the fix window, target instability breaches acceptance limits, or operators cannot establish trustworthy telemetry. Preserve the migration journal and logs.

  1. Stop or isolate client writes and keep automatic reconcilers paused. Confirm rollback authority and target identity.
  2. Restore the protected pre-cutover target using the original evidence directory:
Terminal window
python3 scripts/migrate_postgres.py rollback \
--evidence .tmp/migration-run --confirm-target springmusic
  1. Require backup-checksum validation, a separate pre-rollback dump of the current target, and an original-data fingerprint match before restart.
  2. Check Gateway reachability, application behavior, and the restored target state. Do not assume this state contains the latest source data.
  3. Preserve all post-cutover writes in the pre-rollback dump for explicit reconciliation. They are not merged into the restored data automatically.

Returning users to the VM is a separate decision: confirm source integrity, reconcile any accepted target writes, redirect traffic using the approved procedure, and allow exactly one side to accept writes. Restoring the pre-cutover target alone does not perform these steps.

If cutover failed before a valid backup was recorded, inspect the journal and database with the database owner. Never overwrite the evidence directory or blindly rerun cutover. After a killed process, inspect remaining springmusic-migration-* pods and the stopped Deployment before resuming. The local migration lock does not coordinate different execution hosts.

Transfer configuration, acceptance evidence, dashboards, incident ownership, and source-retention decisions.

Transfer the reviewed configuration revision, workload and Gateway inventory, source manifest, migration journal, backup locations, acceptance results, dashboard URL, and rollback decision. Keep credentials out of the handover document; reference the approved secret store instead.

Agree an initial 24-72 hour stabilization window appropriate to the workload. Assign named incident and database recovery owners, confirm retention and restore procedures, and test alert delivery before relying on it. Flex backups complement migration dumps; a verified dump rollback is not proof of managed-service recovery.

Exit stabilization only with sustained business health, complete telemetry, no unresolved critical issues, and operations sign-off. Resume paused automation deliberately. Keep source data and protected evidence until the agreed retention and reconciliation gates permit decommissioning. Begin HPA and capacity experiments only after stabilization, in a separate change window.

Code & registry github.com Executable migration and recovery workflow Use the reference repository for exact command prerequisites, evidence formats, safeguards, and validation coverage. Open the repository
GOAL

Stabilize and Optimize

Return to the Migration Framework's Optimize loop: collect representative operating evidence, identify the limiting layer, implement one controlled change, and validate reliability, performance, and cost before keeping it.

MigrateOptimizeOverview In 7 trails

Optimize starts when workloads run on STACKIT and real operating data is available. The module converts post-cutover observations into measurable improvements for performance, stability, and cost efficiency.

Optimize is not a one-time task. It is an iterative cycle that can overlap with early stabilization and post-cutover care.

Many right-sizing and tuning decisions are only reliable under real load patterns. After cutover, teams can use production telemetry to separate assumptions from actual behavior.

  1. Collect runtime evidence: utilization, latency, error rates, throughput, and cost drivers.
  2. Identify bottlenecks and waste patterns at workload, platform, and data layers.
  3. Prioritize actions by business impact, risk reduction, and FinOps effect.
  4. Implement tuning changes in controlled increments.
  5. Validate outcomes against SLO, reliability, and cost targets.
  6. Feed lessons learned into future migration waves and operating standards.
  • Rightsizing: Align compute, storage, and network capacity with actual demand profiles.
  • Performance tuning: Improve latency and throughput through configuration, scaling, and architecture adjustments.
  • Reliability hardening: Reduce incident frequency through better resilience, observability, and failure handling.
  • FinOps controls: Improve cost transparency, remove waste, and optimize run-rate efficiency.

Optimization decisions should be based on runtime evidence, not assumptions. For practical implementation, combine workload telemetry, alerting, and controlled infrastructure changes.

  • Managed observability baseline: Use STACKIT Observability to collect metrics, logs, and traces with Grafana, Prometheus, Thanos, Loki, and Tempo.
  • Detection logic: Define explicit thresholds and observation windows for low utilization and overload conditions.
  • Run path: Apply rightsizing through IaC changes (for example VM flavor changes) with rollback checkpoints.
  • Validation loop: Re-measure SLO, error rates, and run-cost after each tuning increment.
Filters

Within a group every tick widens the list. Groups narrow each other.

Framework

Status

Topics

Asset title
Framework
Asset type

For Replatform workloads on Kubernetes, optimization spans multiple layers and should be coordinated as one control loop.

  • Pod scaling: Use HPA to adapt replica count to workload pressure with explicit min/max limits.
  • Node pool scaling: Keep sufficient cluster headroom and tune machine type (flavor) for CPU/memory density requirements.
  • Ingress scaling: Re-evaluate load balancer service plan when ingress throughput or connection behavior becomes a bottleneck.
  • Storage rightsizing: Select storage classes based on performance requirements for persistent workloads.
  • Validation discipline: Re-check latency, error rate, and cost after every incremental tuning change.

Primary inputs

Cutover reports, incident trends, SLO measurements, telemetry baselines, and cost reports.

Optimization outputs

Prioritized improvement backlog, validated tuning changes, and updated runbook standards.

Governance outcome

Clear trade-off decisions between performance, resilience, and cost with documented ownership.

  • Optimize follows technical migration delivery in Migrate.
  • Optimize can run in parallel with early post-cutover care activities, while ownership for this care model is covered in the Run phase.
  • Deeper architectural redesign remains in Refactor.
STACKIT LogoSTACKIT Logo
Optimize Replatform Spring Boot on Kubernetes with Rightsizing and Pod Scaling STACKIT · Runbook Open asset ↗

After migration acceptance and stabilization, use measured workload behavior to choose one optimization at a time. This asset covers pod resources, worker capacity, optional HPA, and PostgreSQL Flex. It does not claim those changes were exercised during the migration test.

Continue the same Terraform, Helm, Spring Music JAR, PostgreSQL Flex databases, and Observability deployment used for provisioning, rehearsal, cutover, and rollback. Do not introduce a second sample or perform capacity experiments during the migration window.

Code & registry github.com Spring Boot Kubernetes Replatform reference Use the same versioned variables, deployment resources, and dashboard as the migration and stabilization workflow. Open the repository

Open grafana_dashboard_url or the SCF Replatform folder. Terraform manages eight panels.

Replatform Grafana dashboard showing cluster CPU and memory, one Spring Boot pod, application requests, and PostgreSQL Flex metrics over one hour

Snapshot from the reference deployment on September 25, 2026, 14:41-15:41 UTC. This is one hour of low-load test operation, not a representative production sizing baseline. Read application activity alongside database availability and pressure before selecting an optimization candidate; the panel interpretations below explain the limits of these signals.

Verify both scrape jobs have actual up=1 samples before interpreting the dashboard. Some cluster panels include fallback values, so a rendered zero is not evidence of zero consumption. Scope queries to the intended cluster and database when a datasource contains multiple workloads. Use additional telemetry and business tests for latency percentiles, errors, and recovery objectives.

Collect a representative baseline that includes busy periods, scheduled work, JVM warmup, and database maintenance. Agree the observation window, business SLOs, capacity headroom, and cost target before making a change. Fourteen days can be a starting observation window, not a rule.

  • Pod-resource candidate: throttling, restart, heap, or working-set pressure isolated to the application or a sidecar.
  • Worker-capacity candidate: pending pods, insufficient allocatable resources, or insufficient rollout headroom across the pool.
  • Database candidate: connection, transaction, lock, storage, or query pressure correlated with business latency.
  • Scale-in candidate: sustained spare capacity after allowing for peaks, rollout, and recovery, with no unresolved critical incidents.

Retain the baseline, previous configuration, rollback plan, and decision thresholds. Missing metrics, failed alert delivery, or synthetic traffic alone are insufficient evidence for production downsizing.

Keep database and application signals in the same review. Increasing pod count increases connection demand and can move the bottleneck to Flex. Separate connection-pool limits, expensive queries, lock contention, and storage pressure from genuine CPU or memory shortages.

Use the PostgreSQL Flex monitoring guidance to interpret service metrics alongside application behavior.

The reference exposes postgres_flex_cpu, postgres_flex_ram, postgres_flex_replicas, postgres_flex_storage_class, and postgres_flex_storage_size. CPU, RAM, and the Single or Replica selection resolve a flavor from the project’s current catalog. Select an offered combination; do not assume arbitrary values or an in-place transition are supported.

Review the plan and service constraints before approval. Treat a database replacement as a new migration with verified recovery, not a routine resize. Storage growth and service-plan transitions may not be reversible by restoring previous variable values. Confirm the recovery path and required maintenance window before changing them.

The tested migration rollback recovers application data; it does not undo infrastructure resizing or prove Flex managed-service restore. Validate the required recovery method separately.

STACKIT documentation docs.stackit.cloud PostgreSQL Flex flavors and performance classes Open the documentation
From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

The reference declares resources in the Spring Boot Deployment in main.tf, not in dedicated CPU or memory variables. Java requests 100m CPU and 512Mi memory, with limits of 500m and 1Gi; JAVA_TOOL_OPTIONS sets a 128 MiB initial and 512 MiB maximum heap. Each of the two exporter sidecars has its own resource budget.

Compare actual working set, heap, non-heap memory, throttling, startup behavior, and sidecar usage before editing the Deployment. Leave room beyond the Java heap for threads and native memory. A resource edit can roll pods and interrupts a single-replica workload; schedule and validate it accordingly. Do not invent unsupported springboot_cpu or memory variable overrides.

Qualify metrics, replica ownership, application safety, and per-pod telemetry before a bounded HPA experiment.

HPA compares observed pod CPU utilization with the configured target and adjusts replicas within minimum and maximum bounds. Its resource metric depends on realistic requests and an available Kubernetes metrics API; Grafana scrape success does not prove that API works. Resource utilization also includes the sidecar budgets. HPA cannot create worker capacity by itself.

Kubernetes metrics APIHPA: CPU target + replica boundsSpring Boot DeploymentWorker capacity + schedulingFlex connection budget observed utilizationadjust replicasmeasure each podschedule within headroomcombined connection demand

Before a multi-replica experiment, review session state, shared writes, initialization, and database connection limits. The current application/exporter scrape uses one load-balanced Service endpoint; replicas can be sampled interchangeably rather than as separate time series. Establish per-pod application scraping and avoid duplicate database aggregation before trusting scaled request rates or totals. These extensions are not part of the validated single-replica path.

Only after migration and rollback operations have finished, test bounded HPA in a separate approved experiment. These illustrative bounds are not production sizing recommendations:

enable_springboot_hpa = true
springboot_hpa_min_replicas = 1
springboot_hpa_max_replicas = 3
springboot_hpa_target_cpu_utilization_percentage = 70

The Deployment also declares springboot_replicas in Terraform. Inspect later plans for competing replica changes and establish an explicit ownership policy before unattended HPA operation. The migration script refuses HPA-managed targets; disable HPA before any later migration or rollback.

Review the plan, then inspect HPA behavior with the configured kubeconfig:

Terminal window
terraform plan -var-file=env.tfvars -out=tfplan.optimize
terraform apply tfplan.optimize
kubectl get hpa,pods -n springboot
kubectl describe hpa springboot -n springboot
kubectl top pods -n springboot --containers

Supply enough worker headroom and account for pool capacity, zone constraints, and rollout disruption.

STACKIT SKE Grafana dashboard showing actual CPU and RAM usage versus requests and limits, one node, 17 running pods, no pending or failed pods, and API server activity

The SKE dashboard shows the same 14:41-15:41 UTC interval on September 25, 2026. Actual CPU usage is about 2%, while CPU requests reserve about 34% of cluster capacity. This difference illustrates why scheduling reservations and measured consumption must be reviewed together. The 17 running pods include platform components, not 17 Spring Boot replicas; the workload dashboard above shows the single application pod. No failed or pending pods at this point is a useful health signal, not proof of peak-load or failure tolerance.

Tune node_pool_minimum, node_pool_maximum, and node_pool_machine_type from aggregate requests, observed demand, system overhead, and rollout headroom. Equal minimum and maximum values fix the pool size; increasing an HPA maximum cannot overcome that capacity limit.

The reference configures one node pool. Additional pools and zone placement require an explicit architecture extension. A node pool’s availability zone cannot be changed in place; a different zone needs a new pool name and a reviewed migration plan. Check actual SKE capacity and planned worker replacement before applying a flavor or topology change.

STACKIT documentation docs.stackit.cloud SKE node-pool management Open the documentation

The implemented entry point is Envoy Gateway with HTTPRoutes, not legacy Ingress. Compare Gateway and service behavior with application and database latency before changing worker size. The optional in-cluster load generator bypasses the public Gateway, DNS, and TLS path; add an approved external test for end-to-end traffic. No measured public-throughput limit is claimed here.

Storage layer (persistent volume performance)

Section titled “Storage layer (persistent volume performance)”

Spring Music stores its authoritative data in Flex. There is no application PersistentVolume to rightsize in this baseline. node_pool_volume_size concerns worker storage, not database capacity. Use the Flex storage controls for album data and review growth, query I/O, retention, and recovery together. Add Kubernetes storage only for a separately designed persistence need.

Optimize workflow for Replatform workloads

Section titled “Optimize workflow for Replatform workloads”
  1. Record representative metrics, business acceptance limits, current configuration, and cost.
  2. Select one hypothesis: pod budget, worker capacity, Gateway, or database pressure.
  3. Specify the expected improvement and rollback threshold; verify required backups and recovery.
  4. Review a saved Terraform plan, reject unrelated changes, and apply in the approved window.
  5. Validate rollout, Gateway, album data, actual scrapes, latency, errors, capacity, and cost against the baseline.
  6. Keep the change only when the agreed observation window meets acceptance; otherwise follow the pre-approved reversal or recovery procedure.

For reversible configuration changes, restore the previous reviewed values and inspect a new plan before applying. Do not assume a smaller database or restored storage class is supported. When HPA was the experiment, disable it and restore the intended replica count through the reviewed configuration; confirm that the Deployment is stable afterward.

Record before/after evidence, configuration revision, business results, and cost impact. Database migration rollback is not a substitute for reversing an optimization change.

The live reference test proved the single-replica workload, data migration and rollback, and dashboard/scrape path. It did not establish autoscaling behavior, optimal sizing, production load capacity, or high availability. Capture fresh evidence for each of those decisions.

External source kubernetes.io Kubernetes Horizontal Pod Autoscaler Review the upstream control-loop behavior, metrics prerequisites, and scaling constraints before enabling autoscaling. Open external site Leads off the trail
Trail historyActive 2 of the last 12 weeksLWUpdatedNo updates · 1 bar = 1 week i
Maintainers
LWLukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwnerActive 10 of the last 12 weeks · 47 updatesSTACKITwww.linkedin.com/in/lukas-weberruß-a360b081Contributed in STACKIT
Show full history (1 more)