Skip to content
Beta

Optimize Rehost Spring Boot with Observability and VM Rightsizing

In 2 trails

Last updated on

This asset continues the same Spring Boot and PostgreSQL reference implementation used for provisioning, migration, cutover, and stabilization. It does not introduce another example or repository. The existing Terraform variables, Ansible configuration, Observability instance, and validation workflows remain the technical baseline for Optimize.

The Optimize extension answers one practical question: how to detect overprovisioning or underprovisioning and then change VM or storage capacity through a controlled IaC workflow.

Code & registry github.com STACKIT CMF Rehost Spring Boot repository Continue with the same Terraform and Ansible reference implementation used by the preceding Rehost migration steps. Open the repository Rehost implementation asset

Use the managed STACKIT Observability stack as evidence source.

  • Prometheus: metric collection
  • Thanos: long-term metric retention
  • Grafana Loki: log analysis
  • Grafana Tempo: distributed traces
  • Grafana: dashboards and visualization

Architecture reference:

Use the dashboard to review infrastructure saturation, application health, request behavior, and alert history together before changing VM capacity.

Grafana dashboard for Rehost Spring Boot observability and VM rightsizing decisions

Define technical thresholds before changing capacity.

  • Candidate for downsizing: CPU p95 under 30% and memory p95 under 50% for at least 14 days.
  • Scale-up candidate: CPU p95 over 75% or memory p95 over 80% during business load windows for at least 3 consecutive days.
  • Stability guardrail: No unresolved critical alerts and no regression in error-rate SLOs.

Keep thresholds workload-specific and validate with business traffic patterns.

Database visibility for optimize decisions

Section titled “Database visibility for optimize decisions”

For stateful Rehost workloads, include database signals in the same dashboard review cycle.

  • Connection pressure: active connection trend and burst behavior.
  • Database growth: database size progression over time.
  • Transaction behavior: commit/rollback trend for stability checks.

Use these metrics together with VM signals to avoid CPU-only or memory-only optimization decisions.

PostgreSQL Flex optimization path for migrated data tiers

Section titled “PostgreSQL Flex optimization path for migrated data tiers”

If the Rehost workload later moves its database tier to PostgreSQL Flex, include database-tier rightsizing and tuning in the same optimize cycle.

Use the offered Flex combinations and storage performance limits to assess a database-tier change. Keep the VM thresholds above workload-specific; the product catalog does not supply acceptance criteria or prove that an infrastructure change is reversible.

From the STACKIT docsFlavors and performance classes › FlavorsSource updated 06.07.2026 · copied 05.10.2026

Notes

  • CPU and RAM is always per node.
  • The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the max_connections limit.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

From the STACKIT docsFlavors and performance classes › Performance ClassesSource updated 06.07.2026 · copied 05.10.2026

Currently, we offer three types of instances. For each type there is a different set of flavors available.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  1. Collect baseline metrics and traces for a representative period.
  2. Confirm optimization candidate with dashboards and alert history.
  3. Plan capacity change and rollback checkpoint.
  4. Apply VM size change with Terraform/OpenTofu.
  5. Re-validate latency, error rates, throughput, and cost.
  6. Keep or revert based on objective acceptance criteria.

Example A: Downsize after sustained low utilization

Section titled “Example A: Downsize after sustained low utilization”

Update VM sizing in env.tfvars:

machine_type = "g3i.2"

Apply and inspect plan output:

Terminal window
terraform plan -var-file=env.tfvars
terraform apply -var-file=env.tfvars

Then validate:

  • Service health (systemctl status, synthetic checks)
  • p95 latency and error-rate trend
  • Cost delta in reporting window

Example B: Scale up under sustained overload

Section titled “Example B: Scale up under sustained overload”

Update VM sizing in env.tfvars:

machine_type = "g3i.4"

Apply and validate with the same post-change checks.

If SLOs regress after rightsizing, roll back by restoring the previous machine_type and re-applying IaC. Treat rollback as a standard runbook step, not as an emergency-only path.

  • Depending on platform constraints and machine type, resize can require restart or replacement. Confirm behavior in terraform plan before apply.

In Rehost scenarios, CPU and memory are only one side of rightsizing. Storage performance can also become the limiting factor.

  • When to investigate storage: elevated I/O wait, unstable latency under write-heavy load, or throughput saturation despite available CPU.
  • What to select: a storage service plan and performance class that matches the observed IOPS and throughput profile.
  • Guidance: Block Storage service plans

Select the performance class before provisioning

Section titled “Select the performance class before provisioning”

A Block Storage performance class defines the maximum IOPS and throughput available to the complete volume. Application, database, operating-system, and backup access share this performance envelope. Select the class from measured peak demand, latency requirements, backup activity, and explicit growth headroom before creating the volume.

From the STACKIT docsService plans › Currently available Service Plans (performance classes)Source updated 22.04.2026 · copied 05.10.2026

The following table lists currently available performance classes for the EU01 region:

IOPS - Input/Output Operations per second

Throughput - Throughput in Megabytes per second

Thus, the classes used can be distinguished in detail based on the naming. Example: “Block Storage Premium - Performance Class 2” corresponds to SSD hard disks with max. 1000 IOPS and max. 100 Mbyte/s throughput.

What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

For the Rehost baseline, the operating system, Spring Boot application, and PostgreSQL data share the boot volume. Changing its performance class therefore requires a controlled replacement target:

  1. Select the new class from observed IOPS, throughput, latency, and I/O-wait data.
  2. Verify backup and database-level rollback readiness.
  3. Provision the replacement VM and boot volume through IaC with the selected class.
  4. Reapply the Ansible configuration and restore or migrate the workload data.
  5. Validate application behavior, data integrity, storage latency, backup coverage, and cost before switching.

Use a separate data volume when storage capacity or performance must evolve independently from the VM lifecycle. To change its performance class, create a new volume in the required availability model with sufficient capacity, stop writes, migrate and verify the data, switch the attachment or mount, and retain the source volume until acceptance and rollback gates have passed.

Migrate data from Block Storage

Treat storage checks as part of the same Optimize loop and validate latency, error behavior, recovery, and cost impact after any change.

Asset historyActive 7 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
Maintainers
TMTobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172LWLukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwnerActive 10 of the last 12 weeks · 47 updatesSTACKITwww.linkedin.com/in/lukas-weberruß-a360b081Contributed in STACKIT
Show full history (7 more)