Purpose
Section titled “Purpose”Use this runbook after a relocated VM has passed cutover acceptance and produced representative telemetry on STACKIT. The objective is to correct initial compute and storage assumptions without trading cost reduction for instability or changing application architecture.
Initial migration sizing and post-cutover rightsizing are separate decisions. Keep the migration profile until target measurements cover normal demand, peak periods, batch processing, backups, and the workload’s known seasonal or month-end events.
Entry criteria
Section titled “Entry criteria”- Cutover is accepted and there are no unresolved migration defects that distort measurements.
- Monitoring, logs, alerts, backups, and application health checks are working on STACKIT.
- The current machine type, volume sizes, performance classes, and initial sizing assumptions are recorded in infrastructure as code or another version-controlled configuration.
- The observation window and workload-specific service-level thresholds are approved before candidate selection.
- A maintenance window, tested recovery point, rollback type, and technical decision owner exist.
Build the target baseline
Section titled “Build the target baseline”Correlate infrastructure and application behavior instead of optimizing from one metric.
- CPU: Review utilization distribution, p95 and peak demand, load, steal time, run queue, and burst duration. Investigate sustained saturation, contention, or unused cores across all representative windows.
- Memory: Review working set, available memory, cache, swap, paging, OOM events, and application heap. Investigate swap or OOM pressure and consistently unused allocation without cache benefit.
- Storage: Review IOPS, throughput, latency, queue depth, I/O wait, block size, backup overlap, and growth. Investigate class saturation, unstable latency, capacity pressure, or unused headroom.
- Network: Review throughput, packet loss, retransmits, connection pressure, and latency. Exclude a network bottleneck before attributing pressure to CPU or storage.
- Application: Review request rate, p95 and p99 latency, errors, job duration, timeouts, and dependency health. Reject a cost improvement that violates workload acceptance thresholds.
Exclude periods affected by migration copying, one-time cache warm-up, failed dependencies, or measurement gaps unless the same condition is expected during normal operation. Preserve excluded periods and the reason for exclusion in the evidence record.
Select a rightsizing action
Section titled “Select a rightsizing action”Classify the finding before changing capacity:
- No change: The current profile meets performance, resilience, and cost expectations.
- Compute downsize: CPU and memory retain approved headroom across representative demand.
- Compute upsize or family change: CPU, memory, or CPU-to-memory ratio constrains the workload.
- Move to a non-overprovisioned type: Sustained CPU demand, latency sensitivity, or steal time requires more predictable CPU access.
- Increase volume capacity: Forecast usable capacity reaches the approved threshold.
- Change storage performance: IOPS or throughput limits, latency, or I/O wait indicate a class mismatch after application and guest causes are excluded.
- Investigate first: The limiting signal is caused by configuration, application behavior, dependency latency, network loss, or insufficient evidence rather than resource capacity.
Change one dominant dimension at a time where practical. This keeps the result attributable and makes rollback decisions defensible.
Change the machine type
Section titled “Change the machine type”- Select the smallest current machine type that meets measured CPU and memory demand plus the approved headroom. Re-evaluate the type family and CPU-overprovisioning decision instead of changing only the size suffix.
- Confirm regional availability, quota, processor architecture, guest support, licensing, maintenance behavior, and the expected operating-system impact.
- Update the version-controlled configuration and inspect the complete infrastructure plan. Stop if it proposes an unintended server, volume, NIC, address, or attachment replacement.
- Capture a tested recovery point, stop the application cleanly, and execute the machine-type resize in the approved maintenance window.
- Verify boot, guest-visible CPU and memory, drivers, disks, NICs, routes, services, monitoring, and application health before restoring traffic.
- Compare application and infrastructure telemetry with the pre-change baseline for the defined validation window.
By changing the machine-type the server will have a short downtime.
Resizing a server
$ stackit server resize <SERVER_ID> --machine-type <TYPE_NAME>
$ stackit server resize xxxxxxxx-xxx-xxxx-xxxx-xxxxxxxxxxxx --machine-type g2i.4<TYPE_NAME>should be replaced with the name of the new machine-type that should be used<SERVER_ID>should be replaced with the ID of the server you want to resize
After the resize it could be necessary to do changes in your operating system.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Use the portal, CLI, API, or infrastructure-as-code workflow owned by the workload, but do not mix control paths without reconciling state afterward.
Increase volume capacity
Section titled “Increase volume capacity”A managed STACKIT volume can only be updated to a larger size. Capacity growth does not by itself change the selected performance class because class limits are independent of volume size.
- Confirm the capacity forecast, backup impact, quota, and the maximum size supported by the guest partition table and file system.
- Update the volume size through the controlled provisioning path and inspect the plan.
- Extend the guest partition, physical volume, logical volume, and file system only as required by the operating-system layout.
- Verify usable capacity, file-system health, backup behavior, monitoring, and application I/O.
Volume shrinking requires a new smaller volume and a data migration. Do not attempt to reduce the managed volume and assume the guest file system will make that operation safe.
Run the following command to resize a volume:
stackit beta volume resize <VOLUME_ID> --size <DRIVE_SIZE> --project-id <PROJECT_ID>Replace the placeholders in the command as follows:
<PROJECT_ID>: Your STACKIT Project ID<VOLUME_ID>: The ID of the Volume you want to resize<DRIVE_SIZE>: The new size of the Volume in gigabytes (GB)
After resizing the volume, you may need to perform additional steps to utilize the extra disk space, depending on your operating system.
These steps typically involve extending the file system to recognize and use the newly allocated space.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Change storage performance
Section titled “Change storage performance”Select the new performance class from measured IOPS and throughput requirements independently. The chosen class must satisfy both limits and include backup, recovery, burst, and growth headroom.
The following table lists currently available performance classes for the EU01 region:
| Performance class | Name | Max. IOPS | Max. Throughput (MB/s) |
|---|---|---|---|
| Performance class 0 | storage_premium_perf0 | 120 | 25 |
| Performance class 1 | storage_premium_perf1 | 500 | 50 |
| Performance class 2 | storage_premium_perf2 | 1000 | 100 |
| Performance class 4 | storage_premium_perf4 | 2000 | 150 |
| Performance class 6 | storage_premium_perf6 | 5000 | 200 |
| Performance class 8 | storage_premium_perf8 | 10000 | 250 |
| Performance class 10 | storage_premium_perf10 | 15000 | 300 |
| Performance class 12 | storage_premium_perf12 | 20000 | 350 |
| Performance class 13 | storage_premium_perf13 | 20000 | 700 |
| Performance class 14 | storage_premium_perf14 | 25000 | 400 |
| Performance class 15 | storage_premium_perf15 | 25000 | 800 |
| Performance class 16 | storage_premium_perf16 | 30000 | 450 |
| Performance class 17 | storage_premium_perf17 | 30000 | 900 |
| Performance class 18 | storage_premium_perf18 | 35000 | 500 |
| Performance class 19 | storage_premium_perf19 | 35000 | 1000 |
| Performance class 20 | storage_premium_perf20 | 40000 | 550 |
| Performance class 21 | storage_premium_perf21 | 40000 | 1100 |
| Performance class 29 | storage_premium_perf29 | 60000 | 1500 |
IOPS - Input/Output Operations per second
Throughput - Throughput in Megabytes per second
Thus, the classes used can be distinguished in detail based on the naming. Example: “Block Storage Premium - Performance Class 2” corresponds to SSD hard disks with max. 1000 IOPS and max. 100 Mbyte/s throughput.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Treat a performance-class change as a replacement workflow unless the current API and provisioning plan explicitly prove an in-place operation for that resource. With Terraform-managed volumes, inspect the plan for replacement before approval.
- Create a target volume with the required availability model, capacity, encryption, and performance class.
- Attach it in a maintenance-safe state and prepare the partitioning, file system, permissions, mount options, and monitoring.
- Copy the bulk data while the workload is online only where application consistency permits it.
- Stop writes, run the final synchronization or application-native consistency procedure, and verify checksums, record counts, or recovery state.
- Shut down the VM before final detach and attach operations, switch the mount or device mapping, then start and validate the workload.
- Retain the previous volume without writers for the approved rollback window and delete it only after backup and acceptance evidence are complete.
For boot volumes or stateful systems that cannot safely move at file level, use a tested snapshot, image, block-copy, or application-native migration procedure. Define the new boot and rollback path before the maintenance window.
Acceptance and rollback
Section titled “Acceptance and rollback”Accept the change only when all workload-specific criteria pass:
- VM and application services start without new warnings or device changes.
- Request latency, error rate, throughput, and batch duration remain within approved limits.
- CPU, memory, storage, and network retain the documented headroom during representative demand.
- Backup, monitoring, alerts, administrative access, and security controls remain functional.
- The measured cost and capacity result matches the expected improvement.
For a failed machine-type change, restore the previous supported type through the same control path and repeat the boot and application checks. For a failed storage migration, stop target writes and restore the previous attachment or data source according to the consistency plan. Preserve all evidence even when the candidate is rejected.
Rightsizing record
Section titled “Rightsizing record”Record the observation period, excluded intervals, metric queries, current and candidate profiles, headroom, expected cost effect, infrastructure plan, maintenance timeline, test results, decision, and rollback outcome. Schedule another review when workload demand, application architecture, retention, or growth assumptions materially change.
Primary references
Section titled “Primary references”Asset historyActive 2 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
- LWLukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwner
Lukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwnerActive 10 of the last 12 weeks · 47 updateswww.linkedin.com/in/lukas-weberruß-a360b081