Create a Nested VMware ESXi Lab on STACKIT
Last updated on
Purpose and boundaries
Section titled “Purpose and boundaries”Build an isolated VMware source environment when a migration lab has no existing ESXi host. A STACKIT Linux server runs QEMU/KVM, which runs ESXi, which runs the test workload VM. This is optional laboratory infrastructure, not a prerequisite of the VMware Relocate Trail. For a real migration, use the customer’s existing, appropriately licensed VMware environment.
The Linux host is sometimes called a carrier: it supplies compute, virtual devices and the local ESXi network. It is neither a Coriolis worker nor the migrated application VM. Distinguish all three layers when reserving memory, diagnosing a failure or shutting down the lab.
Tested profile and operator inputs
Section titled “Tested profile and operator inputs”The reference used the following profile. Recheck the catalog, quotas and CPU exposure on your actual server; a successful test does not guarantee identical behavior after a different placement.
| Layer | Tested configuration | Acceptance requirement |
|---|---|---|
| STACKIT Linux server | Intel c3i.8, 8 vCPU / 16 GiB, eu01-1, Ubuntu 24.04, 80-GiB boot volume. | VMX/EPT exposed, sufficient capacity and reserved Linux/QEMU headroom. |
| Virtualization | QEMU 8.2.2, Q35, KVM required, host CPU with VMX exposed. | Hardware acceleration enabled, with no software fallback. |
| ESXi | ESXi 8.0 Update 3e, build 24677879, 4 vCPU / 10 GiB, BIOS boot. | CPU/memory recognized and installation completed. |
| ESXi storage | Two separate 32-GiB QCOW2 images on Q35 AHCI ports. | System and blank datastore disks independently identified. |
| ESXi networking | VMXNET3; ESXi 10.0.2.15/24, Linux bridge 10.0.2.2/24. | Private management works without a public bridge uplink. |
| First inner guest | Alpine live boot, 1 vCPU / 512 MiB, 2-GiB thin disk, E1000E, hardware version 20. | Real guest execution; not yet application or migration acceptance. |
The earlier minimal KVM probe used c3i.2; it was not the ESXi installation host. Sparse images
still consume real disk capacity as guests write, and the outer host needs memory beyond ESXi’s allocation.
Prepare an approved project and dedicated Linux server through the STACKIT Portal or your reviewed infrastructure workflow. Record server, volume, network, NIC and security-group IDs for cleanup. Restrict SSH to approved operator addresses, use key authentication and verify the host key. Do not expose ESXi or QEMU console ports publicly for installation.
| Input | Meaning and source |
|---|---|
LINUX_HOST | Verified management address of the dedicated STACKIT Linux server. |
ESXI_ISO | Absolute path on that server to the authorized installer ISO. |
ESXI_DIR | New private image/socket directory; this guide uses /var/lib/scf-esxi. |
| Lab addresses | Non-overlapping subnet, Linux bridge address and static ESXi management address. |
| TLS trust | Verified ESXi hostname and certificate/CA, obtained through the console or trusted administration channel. |
Verify hardware virtualization
Section titled “Verify hardware virtualization”Run on the dedicated Ubuntu Linux host, not in ESXi or the documentation container:
set -euo pipefailgrep -qw vmx /proc/cpuinfogrep -qw ept /proc/cpuinfosudo -n modprobe kvm_inteltest -c /dev/kvmtest "$(cat /sys/module/kvm_intel/parameters/nested)" = Ytest "$(cat /sys/module/kvm_intel/parameters/ept)" = Ysudo -n apt-get updatesudo -n apt-get install --yes qemu-system-x86 qemu-utilsqemu-system-x86_64 --versionStop if extensions, nested paging or device access are unavailable. Guest configuration cannot
restore virtualization extensions hidden by the outer platform. Do not reload modules on a shared
host blindly. Require -accel kvm on every launch; ESXi boot and inner-guest execution are separate
gates beyond a generic KVM probe.
Prepare disks and networking
Section titled “Prepare disks and networking”Use the working storage profile
Section titled “Use the working storage profile”The initial NVMe configuration reached the installer but exposed no usable installation disk; the precise cause was not isolated. PVSCSI then failed because ESXi requested MSI-X while the tested QEMU 8.2.2 device model provided MSI. This was a controller-model mismatch, not a proven CPU failure.
Use Q35 AHCI instead. In this machine profile, QEMU’s ide-hd devices on ide.0 and ide.1 attach
to AHCI ports and ESXi binds vmw_ahci. Do not substitute VirtIO or PVSCSI by assumption.
The installation used unmodified media, without injected drivers or unsupported-CPU flags.
Copy authorized media to the host and verify publisher integrity information where available. A locally calculated hash proves copied-byte identity, not media authenticity; the original lab’s local ISO digest was not independently checked against a publisher digest.
ESXI_ISO='/absolute/path/to/authorized-esxi-installer.iso'ESXI_DIR='/var/lib/scf-esxi'test -f "$ESXI_ISO"sha256sum "$ESXI_ISO"sudo -n install -d -m 0700 "$ESXI_DIR"sudo -n test ! -e "$ESXI_DIR/system.qcow2"sudo -n test ! -e "$ESXI_DIR/datastore.qcow2"sudo -n qemu-img create -f qcow2 "$ESXI_DIR/system.qcow2" 32Gsudo -n qemu-img create -f qcow2 "$ESXI_DIR/datastore.qcow2" 32GDo not repeat image creation for a retained installation. Inspect existing disks and ownership;
run qemu-img check only while the images are not open by QEMU. Preserve their identities when restarting.
Replace user networking with bridge/TAP
Section titled “Replace user networking with bridge/TAP”QEMU user-networking port forwarding failed to complete management TCP handshakes in the test: capture showed SYN and repeated SYN-ACK, but no final ACK. Its exact cause was not proven. A Linux-local bridge/TAP resolved the observed connectivity problem.
Inspect interfaces and all routing tables before selecting the example subnet. Include existing VPN routes. Stop on any overlap or pre-existing interface name rather than replacing unknown configuration:
ip -br addressip -4 route show table allip -6 route show table allip link showCreate the isolated network only after that review:
sudo -n ip link add scf-esxi-br type bridgesudo -n ip address add 10.0.2.2/24 dev scf-esxi-brsudo -n ip link set scf-esxi-br upsudo -n ip tuntap add dev scf-esxi-tap mode tapsudo -n ip link set scf-esxi-tap master scf-esxi-brsudo -n ip link set scf-esxi-tap upip -br address show scf-esxi-brDo not bridge the STACKIT uplink, replace default routes or enable global forwarding/NAT for this installation. Initially this is only a host-to-ESXi network: nested workload egress and connectivity from Coriolis workers require a separate, reviewed routing/firewall design.
Install and boot ESXi
Section titled “Install and boot ESXi”The command below uses the tested CPU, NIC and storage models. Run only one process for these images/socket. QMP remains a privileged Unix socket. For interactive installation, this example adds a loopback-only VNC console as a standard alternative to the reference’s QMP screenshots/key input.
sudo -n test ! -e "$ESXI_DIR/qmp.sock"sudo -n qemu-system-x86_64 \ -name scf-esxi-lab -machine q35 -accel kvm -cpu host,vmx=on \ -smp 4,sockets=1,cores=4,threads=1 -m 10240 \ -no-user-config -nodefaults -vga std -vnc 127.0.0.1:0 \ -qmp "unix:$ESXI_DIR/qmp.sock,server=on,wait=off" \ -serial "file:$ESXI_DIR/serial.log" \ -drive "file=$ESXI_DIR/system.qcow2,if=none,id=system,format=qcow2" \ -device ide-hd,bus=ide.0,drive=system,serial=SCFESXISYSTEM \ -drive "file=$ESXI_DIR/datastore.qcow2,if=none,id=datastore,format=qcow2" \ -device ide-hd,bus=ide.1,drive=datastore,serial=SCFESXIDATASTORE \ -cdrom "$ESXI_ISO" -boot order=d,menu=off \ -netdev tap,id=management,ifname=scf-esxi-tap,script=no,downscript=no \ -device vmxnet3,netdev=management,mac=52:54:00:12:34:56Use a different MAC for additional lab hosts. On the operator workstation, tunnel VNC through
verified SSH and connect a VNC viewer to 127.0.0.1:5900:
LINUX_HOST='your-verified-linux-host'ssh -o StrictHostKeyChecking=yes -N \ -L 127.0.0.1:5900:127.0.0.1:5900 "ubuntu@$LINUX_HOST"- Wait for the installer. If it appears stalled, inspect logs with Alt+F12 and return with Alt+F2.
The last
Starting service vmtoolsdbanner alone was not proof of a hang. - Review and accept the EULA yourself. Verify CPU/memory, VMXNET3 and both disk identities.
- Install only to
SCFESXISYSTEM; preserveSCFESXIDATASTORE. Set the root credential privately in the console, never in command arguments, shared screenshots or source control. - Complete installation. The reference showed a legacy-BIOS advisory, not a fatal CPU check.
- Shut down ESXi cleanly, then relaunch the same images without the ISO: replace
-cdrom "$ESXI_ISO" -boot order=d,menu=offwith-boot order=c,menu=off. - In the Direct Console User Interface (DCUI), configure VMXNET3 management with the example’s
static
10.0.2.15/24address and apply the changes. Verify connectivity to10.0.2.2.
The original transient launch had a one-hour runtime limit. Expiry stopped QEMU, not STACKIT billing, and was not a recorded VMkernel crash. Do not retain an installed host under a disposable test timeout. A stale socket needs process/ownership reconciliation, not blind deletion.
Validate management and the datastore
Section titled “Validate management and the datastore”Use the ESXi console or temporarily enabled, restricted ESXi Shell to inspect the actual NIC, management IP and virtualization capability:
localcli network nic listesxcli network ip interface ipv4 getesxcli hardware cpu global getThe tested NIC used nvmxnet3 with link up; HV Support was 3. The API’s nestedHVSupported
field describes another virtualization-exposure capability and is not the same test.
From the workstation, tunnel management HTTPS to the now carrier-reachable ESXi address:
ssh -o StrictHostKeyChecking=yes -N \ -L 127.0.0.1:18443:10.0.2.15:443 "ubuntu@$LINUX_HOST"Use the ESXi certificate hostname, matching local name resolution and verified certificate/CA trust when opening the Host Client through that tunnel. Obtain or replace certificates through supported administration when necessary; do not accept browser warnings or disable TLS validation as the management design. An HTTPS reverse proxy also needs verified upstream TLS and does not relay VMware disk export on TCP 902.
Create VMFS6 on the verified blank SCFESXIDATASTORE disk, never the installed system disk.
Prefer Host Client storage administration where permitted. For an authorized local shell procedure,
follow Broadcom’s disk identification, partedUtil and vmkfstools -C vmfs6 instructions below.
Inspect before changing partitions:
esxcli storage core device listesxcli storage filesystem listThe tested 32-GiB datastore used a GPT VMFS partition starting at sector 2048 and ending at its inspected last usable sector. Derive those values from the actual disk; do not paste geometry from another device. Confirm the new datastore is mounted and has expected capacity. Local shell administration is not a substitute for the license rights needed by migration APIs.
Boot an inner guest and qualify migration access
Section titled “Boot an inner guest and qualify migration access”Create a small Linux VM through Host Client with the profile above, attach legitimate guest media
and verify an actual console boot. Prefer normal VM creation over a minimal handwritten VMX:
the initial VMX lacked PCIe root ports and failed with No PCIe slot available for Ethernet0.
Standard PCI bridge/root-port entries resolved that configuration error, not a nested-CPU change.
A live Alpine login prompt proves guest execution only. For a migration test, install the guest to disk, configure networking and VMware Tools, and establish application/data baselines. The later Spring Boot/PostgreSQL source was a separate installed Ubuntu VM, not this live-boot probe.
Before using ESXi as a Coriolis source:
- Verify the installed license, not the installer’s generic evaluation message. The initial
Free Hypervisor edition returned
RestrictedVersionfor API datastore creation. Obtain appropriate VMware API/snapshot/CBT/export rights; the reference later used a licensed source, not a restriction bypass. - Create a dedicated migration account and test required operations with that account.
- Check management HTTPS and NFC/NBD TCP 902 from the actual Coriolis worker namespace. Review DNS/TLS, narrow routes/firewall rules and return paths; browser access alone is insufficient.
- Verify exact VM inventory/identity, API snapshot creation/removal, CBT and disk export. Resolve installed-provider compatibility through vendor support before migration.
Troubleshooting and retention
Section titled “Troubleshooting and retention”| Symptom | Evidence-based action |
|---|---|
| KVM cannot initialize | Check VMX/EPT, module state, /dev/kvm and permissions. Do not fall back to TCG. |
| Installer has no disks | Use the tested Q35 AHCI attachment; verify serials rather than weakening checks. |
| PVSCSI rejects MSI-X | Check the version-specific device model; AHCI avoided the tested MSI/MSI-X mismatch. |
| Boot seems stuck | Inspect VMkernel logs and the outer process/runtime limit before declaring a hang. |
| HTTPS handshakes stall | Inspect packets at both ends; bridge/TAP replaced the failing user-networking path. |
| Inner VM fails at Ethernet startup | Inspect PCI bridge/root ports or create a normal Host Client VM definition, preserving disks. |
API reports RestrictedVersion | Verify the actual license; shell/UI success does not establish API migration rights. |
| UI works but disk export fails | Test serving-ESXi TCP 902 and endpoint hostname routing separately. |
For retained operation, persist the bridge/TAP and installed-disk QEMU launch in reviewed network and systemd configuration. Networking must exist before QEMU starts. Keep the same device/disk identities, omit the installer ISO, keep control sockets private and review any restart policy. Back up QCOW2 files with a consistency-aware method, not arbitrary copies of open images.
Hand over cloud IDs, ESXi/QEMU versions, license, CPU exposure, disk serials, datastore, network, certificate/SSH trust and the inner-guest test. Shutdown order is workloads, normal ESXi shutdown, then the outer QEMU/Linux host. Retention incurs costs. Delete only approved lab-owned cloud and optional DNS/publication resources; stopping QEMU does not clean up the STACKIT server or volumes.
Primary references
Section titled “Primary references”Asset historyAdded Oct 4, 2026LWUpdatedNo updates · 1 bar = 1 week i
- LWLukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwner
Lukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwnerActive 10 of the last 12 weeks · 47 updateswww.linkedin.com/in/lukas-weberruß-a360b081