Large File Migration with NFS and fpsync
Last updated on
Use case
Section titled “Use case”- Category: Large file data migration
- Target: STACKIT File Service
- Method: Migration VM with source and target NFS mounts, synchronized with fpsync
Network and architecture constraints (validated)
Section titled “Network and architecture constraints (validated)”- SFS is SNA-attached: STACKIT File Storage is connected to a STACKIT Network Area (SNA), uses private interconnection subnets, and requires SNA routing tables. See File Storage concepts and SNA routing tables .
- Mount access is network-restricted: NFS mounts are controlled by Resource Pool IP ACLs and Share Export Policies, so the migration host IP must be explicitly allowed. See Mounting resource pools and shares .
- SNA is regional: Private SNA traffic is regional; cross-region traffic is possible through public internet. See Network Area concepts .
- Provider comparison is consistent: Comparable NFS services such as AWS EFS mount targets are also private by design (no public IP on mount targets), so a private path is required there as well. See AWS EFS network access .
A STACKIT File Storage is always connected to an STACKIT Network Area (SNA). Therefore, a dedicated subnet of the SNA-Network is reserved and used to interconnect the whole SNA with the STACKIT File Storage (routed per default). It can be restricted by policies, if required. STACKIT File Storage requires routing tables to be enabled in the STACKIT Network Area. More information on how to do so can be found in the routing tables docs.
When configuring the SNA, please ensure the following requirements are met:
- There are sufficient IP networks available within the SNA.
- The minimum size for any new subnet must be set to at least /29.
- For the SFS interconnection, at least one subnet with /28 and four subnets with /29 are required.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Mandatory prerequisites
Section titled “Mandatory prerequisites”- Migration host placement: The migration VM runs in an SNA project and region that can mount the target SFS share.
- Dual mount feasibility: Source and target exports can both be mounted on one migration VM.
- Private L3 path to source NFS: The migration VM can reach the source NFS endpoint over private routing (same private domain, peering, or site-to-site VPN).
- NFS policy alignment: ACL/security policy allows required client IP ranges and NFS traffic in both directions.
- Network feasibility: Latency and throughput are sufficient for parallel sync.
- Permission model aligned: UID/GID and ACL translation rules are defined.
- Consistency model defined: Initial sync, delta sync windows, and final freeze are agreed.
Check the effective share permissions in addition to private reachability. The migration host needs read-write access to the target, and ownership mapping must be validated before the first fpsync run.
When a Share is created, you can optionally pass it a Share Export Policy, to control which IPs can mount the Share, and with which permissions. If you don’t attach any Share Export Policy to the Share, mounting the Share inherits the rules of the Resource Pool. In other words, the IP ACL of the Resource Pool is applied and the client can only mount the Share in read-only mode.
A Share has a field called Mount Path, that looks like this: 10.2.1.1:/rp\_VKL20Ub/my-share. It is mountable the same way as the Resource Pool.
Keep in mind that:
- In order to have read-write access on a Share, you need to create a Share Export Policy beforehand and attach it to the Share.
- A Share does not have a fixed size. By default, every Share in a Resource Pool have access to all the space of the Resource Pool. You can limit the space a Share consumes.
- If you apply a Share Export Policy to the Share, you can define a subset of the network in the Resource Pool IP ACL. If you define a network that is bigger than the Resource Pool IP ACL, then the Resource Pool IP ACL will take precedence.
- To ensure proper ownership, you have to adjust the NFSv4 ID domain to
stackit.cloudbeforehand. This can be done in/etc/idmapd.conffollowed by the bash commandnfsidmap -c.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Not suitable when
Section titled “Not suitable when”- No private path exists: Source NFS is not reachable from the SNA-connected migration host through controlled private connectivity.
- Dual NFS mount is blocked: Policy or network constraints prevent simultaneous mounts.
- Protocol mismatch exists: Source endpoint does not provide compatible NFS access.
VPN feasibility
Section titled “VPN feasibility”- VPN is a valid option: This runbook works with VPN when the VPN connects the source network to the target-side SNA and routes are advertised correctly.
- VPN scope is site-to-site: STACKIT VPN is designed for site-to-site connections and SNA-based projects. See STACKIT VPN product overview .
- Operational readiness required: Validate tunnel stability, MTU behavior, and sustained throughput before cutover.
When to choose this variant
Section titled “When to choose this variant”- Choose NFS+fpsync when: One migration host can mount both source NFS and target SFS directly and you need fast iterative delta cycles.
- Do not choose NFS+fpsync when: Source access is not NFS-compatible or no controlled private connectivity to source can be established.
- Alternative: Use the rclone bridge runbook when dual NFS mounting is not feasible.
Migration VM placement and trade-offs
Section titled “Migration VM placement and trade-offs”- Preferred placement (STACKIT side): Run the migration VM in STACKIT, attached to the SNA. This keeps the SFS write path private and makes target-side routing and ACL control more direct.
- Alternative placement (source side): Run a transfer host near the source only when required by source constraints. This is usually harder for dual-mount operation to SFS and often shifts complexity to relay patterns.
- VPN implication: For source reachability, site-to-site VPN is the preferred helper. In this topology, source traffic to the migration VM runs through the VPN tunnel.
Recommended topology: STACKIT migration VM with site-to-site VPN to source
Section titled “Recommended topology: STACKIT migration VM with site-to-site VPN to source”This topology is the primary pattern for VPN-supported dual mounts in this runbook.
Alternative topology with source-side fpsync client and HAProxy TCP proxy
Section titled “Alternative topology with source-side fpsync client and HAProxy TCP proxy”Use this variant when you want to keep the fpsync client on the source side and avoid mounting NFS on the STACKIT jump host itself.
- Pattern: HAProxy on the jump host works as a TCP pass-through endpoint for NFS traffic (
tcp/2049) towards SFS. - Ingress path: The source-side client reaches HAProxy through the public IP of the jump host.
- Security controls on jump host: Security group / ACL on the jump host allows inbound
tcp/2049only from approved source client public IP ranges. - No target mount on jump host: The jump host forwards NFS sessions and does not need to mount the SFS share locally.
- Shared client model: The source-side migration client can mount source NFS directly and target NFS through the HAProxy endpoint, then run
fpsyncbetween both mount points. - Policy prerequisite: SFS ACL/export policy must allow the effective client path and source IP model of this setup.
Why this can be useful
Section titled “Why this can be useful”- Logging: HAProxy TCP logs provide one central trace point for NFS session attempts, connection errors, and backend availability.
- Timeout control: HAProxy timeout settings (
timeout connect,timeout client,timeout server) give explicit control over stuck or long-running connections. - Operational guardrails: You can apply controlled connection handling and clear failure behavior at a single ingress point.
Trade-offs and risks
Section titled “Trade-offs and risks”- Extra hop: Adds one network hop and one additional component in the data path.
- Proxy bottleneck risk: Jump host sizing and HAProxy tuning become throughput-critical.
- Service semantics validation: NFS over TCP proxying must be validated end-to-end in your target policy and support model.
- High availability needed: Without HA design, the proxy host can become a single point of failure.
Variant comparison: VPN dual-mount VM vs HAProxy TCP proxy
Section titled “Variant comparison: VPN dual-mount VM vs HAProxy TCP proxy”| Criterion | Variant A: STACKIT migration VM with dual mounts (VPN to source) | Variant B: Source client + HAProxy TCP proxy |
|---|---|---|
| Security surface | Smaller runtime chain; fewer middle components in data path. | Extra proxy layer to harden and operate; centralized ingress control possible. |
| Performance | Usually higher throughput potential (direct dual mount, fewer hops). | Additional hop and proxy processing reduce peak throughput in most environments. |
| Stability | Fewer moving parts; depends on VPN and migration VM stability. | Additional failure domain (HAProxy host/service); needs HA design for robustness. |
| Monitoring and logging | Relies mainly on host, NFS client, and network metrics. | Strong centralized TCP visibility and timeout observability at proxy layer. |
| Timeout handling | Mostly OS/NFS client behavior on migration VM. | Explicit timeout control in HAProxy plus client-side timeout behavior. |
| Operational complexity | Simpler baseline architecture. | More components to configure, tune, and troubleshoot. |
Shared end-state for both variants
Section titled “Shared end-state for both variants”For both architectures, the effective migration flow can end at the same fpsync model:
- Source-side or STACKIT-side client has two NFS mount points (source and target).
fpsyncruns iterative sync cycles between these mount points.- Final freeze window and final pass remain identical in principle.
Operational flow on the migration VM
Section titled “Operational flow on the migration VM”- Step 1 (source mount): The migration VM mounts the source NFS export through the private path of the selected topology (for this runbook: site-to-site VPN tunnel to source).
- Step 2 (target mount): The same migration VM mounts the SFS target path by NFSv4.1 (TCP 2049) through SNA routing tables.
- Step 3 (sync cycles):
fpsyncruns iterative delta cycles between both mounted paths until cutover. - Step 4 (final pass): After source freeze, run the final
fpsyncpass and complete integrity checks.
Throughput and concurrency tuning
Section titled “Throughput and concurrency tuning”- fpsync worker parallelism: Increase parallel workers with
fpsync -n <parallelism>and tune in controlled increments. - Starting point and scaling: Start with moderate parallelism, observe throughput and error rate, then increase until gains flatten or retries rise.
- NFS client tuning: Validate mount options such as
rsize,wsize, andnconnect(where supported) for both source and target mounts. - Network path quality: Keep MTU, packet loss, and latency stable across the source path (including VPN) and the SNA target path.
- VM sizing: Ensure migration VM CPU, memory, and NIC bandwidth are sufficient for concurrent file traversal and transfer.
- Storage-side limits: Check source export limits and SFS-side throughput behavior so worker scaling does not exceed service-side bottlenecks.
- Workload profile split: Test large-file and small-file datasets in dedicated test sets; small files often require higher metadata parallelism, not only bandwidth.
- Measurement discipline: Track effective MB/s, files/s, retransmits, retries, and server load per tuning step.
Implementation template
Section titled “Implementation template”Phase 1: Prepare migration VM
Section titled “Phase 1: Prepare migration VM”- Harden migration VM and configure logging.
- Mount source and target NFS paths with verified options.
- Run baseline read/write probes and record throughput.
Phase 2: Initial and delta sync
Section titled “Phase 2: Initial and delta sync”- Run initial fpsync pass.
- Capture transfer statistics and error files.
- Schedule delta sync cycles until cutover window.
Phase 3: Final sync and handover
Section titled “Phase 3: Final sync and handover”- Activate final write freeze on source window.
- Run final fpsync pass and verify integrity sample.
- Hand over mounted target path to consuming workload.
Validation checklist
Section titled “Validation checklist”- File integrity: Sample checksum and count verification completed.
- Permission integrity: ACL and ownership spot checks completed.
- Performance evidence: Effective throughput and total duration documented.
- Operational handover: Monitoring and ownership confirmed.
Asset historyActive 5 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
- LWLukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwner
Lukas WeberrußHead of STACKIT Cloud Migration Framework · STACKITOwnerActive 10 of the last 12 weeks · 47 updateswww.linkedin.com/in/lukas-weberruß-a360b081