---
title: Optimize Rehost Spring Boot with Observability and VM Rightsizing
description: 'Practical Optimize runbook for a rehosted Spring Boot workload on STACKIT using managed Observability signals and controlled VM rightsizing driven by IaC.'
sidebar:
  badge:
    text: "STACKIT"
    variant: success
scfAsset:
  managed: false
  category: "runbook"
  external: false
  tags: ["migrate", "optimize", "rehost", "observability", "rightsizing", "spring-boot"]
  maintainers:
    - user: "lukas.weberruss"
      role: true
      website: true
source_url: "https://framework.stackit.cloud/migration/assetcontainer/stackit/optimize-rehosted-spring-boot-observability-rightsizing/"
source_file: "docs/migration/assetcontainer/stackit/optimize-rehosted-spring-boot-observability-rightsizing.mdx"
---

## Continue the reference implementation

This asset continues the same Spring Boot and PostgreSQL reference implementation used for
provisioning, migration, cutover, and stabilization. It does not introduce another example or
repository. The existing Terraform variables, Ansible configuration, Observability instance, and
validation workflows remain the technical baseline for Optimize.

The Optimize extension answers one practical question: how to detect overprovisioning or
underprovisioning and then change VM or storage capacity through a controlled IaC workflow.

<LinkCard
  title="STACKIT CMF Rehost Spring Boot repository"
  description="Continue with the same Terraform and Ansible reference implementation used by the preceding Rehost migration steps."
  href="https://github.com/stackitcloud/stackit-cmf-Rehost-springboot"
/>

<LinkChip href="/migration/assetcontainer/stackit/rehost-automation-spring-boot-terraform-ansible/">Rehost implementation asset</LinkChip>

## Observability setup used for optimization

Use the managed <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">STACKIT Observability</LinkChip> stack as evidence source.

- **Prometheus**: metric collection
- **Thanos**: long-term metric retention
- **Grafana Loki**: log analysis
- **Grafana Tempo**: distributed traces
- **Grafana**: dashboards and visualization

Architecture reference:

- <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/basics/architecture-of-observability/">Architecture of Observability</LinkChip>
- <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/getting-started/create-your-first-observability-service/">Create and configure your first Observability service</LinkChip>

## Observability decision dashboard

Use the dashboard to review infrastructure saturation, application health, request behavior, and
alert history together before changing VM capacity.

<Image
  src={grafanaDashboardScreenshot}
  alt="Grafana dashboard for Rehost Spring Boot observability and VM rightsizing decisions"
  style={{ width: "100%", maxWidth: "1540px", height: "auto", display: "block" }}
/>

## Signals and thresholds for rightsizing

Define technical thresholds before changing capacity.

- **Candidate for downsizing**: CPU p95 under 30% and memory p95 under 50% for at least 14 days.
- **Scale-up candidate**: CPU p95 over 75% or memory p95 over 80% during business load windows for at least 3 consecutive days.
- **Stability guardrail**: No unresolved critical alerts and no regression in error-rate SLOs.

Keep thresholds workload-specific and validate with business traffic patterns.

## Database visibility for optimize decisions

For stateful Rehost workloads, include database signals in the same dashboard review cycle.

- **Connection pressure**: active connection trend and burst behavior.
- **Database growth**: database size progression over time.
- **Transaction behavior**: commit/rollback trend for stability checks.

Use these metrics together with VM signals to avoid CPU-only or memory-only optimization decisions.

## PostgreSQL Flex optimization path for migrated data tiers

If the Rehost workload later moves its database tier to PostgreSQL Flex, include database-tier rightsizing and tuning in the same optimize cycle.

- **Plan the target DB profile**: use <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/basics/plan-your-postgresql-flex-instance/">Plan your PostgreSQL Flex instance</LinkChip> to align expected data growth and concurrency.
- **Tune flavor and storage class**: use <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/">Flavors and performance classes of PostgreSQL Flex</LinkChip> to adjust CPU, RAM, and storage performance.
- **Use the integrated dashboard first**: review <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/dashboard-metrics-in-postgresql-flex/">Dashboard metrics in PostgreSQL Flex</LinkChip> for fast checks.
- **Use scraped observability metrics for deep analysis**: follow <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/observability-metrics-in-postgresql-flex/">Observability metrics in PostgreSQL Flex</LinkChip> for alert and threshold design.

Use the offered Flex combinations and storage performance limits to assess a database-tier change.
Keep the VM thresholds above workload-specific; the product catalog does not supply acceptance
criteria or prove that an infrastructure change is reversible.

> From the STACKIT docs: [Flavors and performance classes › Flavors](https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/#flavors) (Source updated 06.07.2026, copied 05.10.2026)

| Description | ID | CPU | RAM | `max_connections` | `shared_buffers` | `work_mem` | `maintenance_work_mem` | `effective_cache_size` |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Small, Compute optimized | 2.4 | 2 | 4 GB | 95 | 950 MB | 14 MB | 380 MB | 2660 MB |
| Small, Memory optimized | 2.16 | 2 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| Medium, Compute optimized | 4.8 | 4 | 8 GB | 195 | 1950 MB | 14 MB | 780 MB | 5460 MB |
| Medium, Memory optimized | 4.32 | 4 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| Large, Processor optimized | 8.16 | 8 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| X-Large, Compute optimized | 16.32 | 16 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| X-Large, Memory optimized | 16.128 | 16 | 128 GB | 3170 | 31950 MB | 14 MB | 12780 MB | 89460 MB |

### Notes

- CPU and RAM is always per node.
- The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the `max_connections` limit.

> From the STACKIT docs: [Flavors and performance classes › Performance Classes](https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/#performance-classes) (Source updated 06.07.2026, copied 05.10.2026)

| Description | ID | Max. IOPS | Max. throughput (MB/s) |
| --- | --- | --- | --- |
| Performance class 2 | `premium-perf2-stackit` | 1000 | 100 |
| Performance class 4 | `premium-perf4-stackit` | 2000 | 150 |
| Performance class 6 | `premium-perf6-stackit` | 5000 | 200 |
| Performance class 8 | `premium-perf8-stackit` | 10000 | 250 |
| Performance class 10 | `premium-perf10-stackit` | 15000 | 300 |
| Performance class 12 | `premium-perf12-stackit` | 20000 | 350 |

Currently, we offer three types of instances. For each type there is a different set of flavors available.

## Rightsizing workflow

1. Collect baseline metrics and traces for a representative period.
2. Confirm optimization candidate with dashboards and alert history.
3. Plan capacity change and rollback checkpoint.
4. Apply VM size change with Terraform/OpenTofu.
5. Re-validate latency, error rates, throughput, and cost.
6. Keep or revert based on objective acceptance criteria.

## Technical implementation example

### Example A: Downsize after sustained low utilization

Update VM sizing in `env.tfvars`:

```hcl
machine_type = "g3i.2"
```

Apply and inspect plan output:

```bash
terraform plan -var-file=env.tfvars
terraform apply -var-file=env.tfvars
```

Then validate:

- Service health (`systemctl status`, synthetic checks)
- p95 latency and error-rate trend
- Cost delta in reporting window

### Example B: Scale up under sustained overload

Update VM sizing in `env.tfvars`:

```hcl
machine_type = "g3i.4"
```

Apply and validate with the same post-change checks.

## Rollback pattern

If SLOs regress after rightsizing, roll back by restoring the previous `machine_type` and re-applying IaC.
Treat rollback as a standard runbook step, not as an emergency-only path.

## Notes

- Depending on platform constraints and machine type, resize can require restart or replacement. Confirm behavior in `terraform plan` before apply.

## Storage as optimization dimension

In Rehost scenarios, CPU and memory are only one side of rightsizing. Storage performance can also become the limiting factor.

- **When to investigate storage**: elevated I/O wait, unstable latency under write-heavy load, or throughput saturation despite available CPU.
- **What to select**: a storage service plan and performance class that matches the observed IOPS and throughput profile.
- **Guidance**: <LinkChip href="https://docs.stackit.cloud/products/storage/block-storage/basics/service-plans/">Block Storage service plans</LinkChip>

### Select the performance class before provisioning

A Block Storage performance class defines the maximum IOPS and throughput available to the complete
volume. Application, database, operating-system, and backup access share this performance envelope.
Select the class from measured peak demand, latency requirements, backup activity, and explicit
growth headroom before creating the volume.

> From the STACKIT docs: [Service plans › Currently available Service Plans (performance classes)](https://docs.stackit.cloud/products/storage/block-storage/basics/service-plans/#currently-available-service-plans-performance-classes) (Source updated 22.04.2026, copied 05.10.2026)

The following table lists currently available performance classes for the EU01 region:

| Performance class | Name | Max. IOPS | Max. Throughput (MB/s) |
| --- | --- | --- | --- |
| Performance class 0 | storage_premium_perf0 | 120 | 25 |
| Performance class 1 | storage_premium_perf1 | 500 | 50 |
| Performance class 2 | storage_premium_perf2 | 1000 | 100 |
| Performance class 4 | storage_premium_perf4 | 2000 | 150 |
| Performance class 6 | storage_premium_perf6 | 5000 | 200 |
| Performance class 8 | storage_premium_perf8 | 10000 | 250 |
| Performance class 10 | storage_premium_perf10 | 15000 | 300 |
| Performance class 12 | storage_premium_perf12 | 20000 | 350 |
| Performance class 13 | storage_premium_perf13 | 20000 | 700 |
| Performance class 14 | storage_premium_perf14 | 25000 | 400 |
| Performance class 15 | storage_premium_perf15 | 25000 | 800 |
| Performance class 16 | storage_premium_perf16 | 30000 | 450 |
| Performance class 17 | storage_premium_perf17 | 30000 | 900 |
| Performance class 18 | storage_premium_perf18 | 35000 | 500 |
| Performance class 19 | storage_premium_perf19 | 35000 | 1000 |
| Performance class 20 | storage_premium_perf20 | 40000 | 550 |
| Performance class 21 | storage_premium_perf21 | 40000 | 1100 |
| Performance class 29 | storage_premium_perf29 | 60000 | 1500 |

IOPS - Input/Output Operations per second

Throughput - Throughput in Megabytes per second

Thus, the classes used can be distinguished in detail based on the naming. Example: “Block Storage Premium - Performance Class 2” corresponds to SSD hard disks with max. 1000 IOPS and max. 100 Mbyte/s throughput.

For the Rehost baseline, the operating system, Spring Boot application, and PostgreSQL data share the
boot volume. Changing its performance class therefore requires a controlled replacement target:

1. Select the new class from observed IOPS, throughput, latency, and I/O-wait data.
2. Verify backup and database-level rollback readiness.
3. Provision the replacement VM and boot volume through IaC with the selected class.
4. Reapply the Ansible configuration and restore or migrate the workload data.
5. Validate application behavior, data integrity, storage latency, backup coverage, and cost before switching.

Use a separate data volume when storage capacity or performance must evolve independently from the
VM lifecycle. To change its performance class, create a new volume in the required availability
model with sufficient capacity, stop writes, migrate and verify the data, switch the attachment or
mount, and retain the source volume until acceptance and rollback gates have passed.

<LinkChip href="https://docs.stackit.cloud/products/storage/block-storage/how-tos/migrate-data-from-a-block-storage/">Migrate data from Block Storage</LinkChip>

Treat storage checks as part of the same Optimize loop and validate latency, error behavior, recovery,
and cost impact after any change.
