Skip to content
Beta

Migration Framework Expedition: Discovery to Day-2 Operations

Last updated on

TM LogoTM Logo
TM Solutions Corp. Inc.

Migration Framework Expedition: Discovery to Day-2 Operations

Guided end-to-end expedition through the STACKIT Migration Framework, from discovery and R-strategy decisions to factory waves, cutover, and day-2 operations.

PLAN

Capture the Discovery Evidence Base

Every migration decision downstream is only as good as the discovery input. Start by consolidating inventory, dependencies, utilization, and operational constraints into a migration-ready evidence base.

Design and mobilizeDiscoveryOverview In 4 trails

Discovery is one of the first and most critical modules in the Design and Mobilize phase. It refines Rapid Discovery results and adds the depth needed to make architecture and migration-wave decisions with confidence.

The primary objective is to establish a realistic, evidence-based understanding of the current IT landscape, business priorities, and organizational readiness before detailed target design and migration planning are finalized.

Complete baseline

Create a reliable application and infrastructure baseline that goes beyond pure quantities.

Dependency transparency

Identify technical and process dependencies to avoid hidden migration blockers.

Business alignment

Link technical findings with business criticality, timelines, and risk tolerance.

Planning readiness

Produce decision-ready input for target design and migration-wave planning.

Inventory

Comprehensive capture of servers, virtual machines, databases, middleware, and applications.

Dependency analysis

Mapping of communication paths and runtime dependencies between systems and applications.

Resource utilization

Analysis of actual CPU, memory, storage, and I/O behavior over a representative period.

Operational context

Collection of backup, patching, SLA, compliance, and operational constraints.

Application owner input

Structured questionnaires and interviews to validate assumptions and close data gaps.

In practice, Discovery is often run together with STACKIT partners. Partners typically use their own tooling landscape to collect and normalize technical data into a central repository. Many programs also trigger targeted questionnaires for application owners directly from these tools to enrich technical findings with business and operational context.

This combined model improves speed and consistency while keeping stakeholder validation built into the process.

Two Evidence Streams: Technical vs. Human-Driven

Section titled “Two Evidence Streams: Technical vs. Human-Driven”

Discovery intentionally combines two evidence streams that complement each other:

  • Technically derived evidence: Tooling-generated findings from inventory exports, runtime metrics, and dependency signals. This stream provides scale, consistency, and repeatability.
  • Application-owner enrichment (human-driven): Validated business criticality, lifecycle intent, release constraints, and operational realities from owner interviews and questionnaires.

Neither stream is sufficient on its own. Technical evidence without owner context can misclassify critical workloads, while human input without technical grounding can hide coupling and capacity risks. Discovery quality depends on reconciling both streams into one decision-ready view.

The following diagram shows how Discovery transforms technical and stakeholder input into decision-ready outputs for the downstream modules.

Swipe sideways to see the whole diagram
Discovery source-to-decision flow Discovery separates technical and human-driven inputs, runs technical-first and human-enriched analyses, and hands over both insight streams to follow-on modules. Discovery inputsTechnical and automated discoveryInfrastructure inventoryCMDB, VM, database, middleware, storageRuntime and utilization dataCPU, memory, I/O, network and seasonalityIntegration and flow signalsNetwork paths, APIs, identity, data movementAssessment-driven human inputSecurity and compliance contextData classes, controls, audit requirementsOwner and business inputCriticality, release windows, lifecycle plansDiscovery analysis toolingTechnical-first analysesNormalize and correlateUnify records and technical identitiesDependency mappingInfer communication and couplingUtilization and sizing analysisEstimate baseline demand corridorsPreliminary segmentationCluster by stack and environmentHuman-driven analyses (application owner input)Criticality and risk calibrationValidate business impact and constraintsWave feasibility and sequencingReconcile dependencies with release windowsAssumption and gap registerTrack open points and confidenceHandover outputsTool-derived outputsDesignTarget architecture options and sizing factsLanding zonePlatform guardrails and account structure needsMigration planWave backlog, sequencing, and cutover windowsAssessment-validated outputsSecurity and complianceControl needs, data classes, remediation pointsOperating modelRole model, ownership boundaries, process impactBusiness caseValue/risk profile and modernization priorities

Typical Analysis Patterns in Discovery Tooling

Section titled “Typical Analysis Patterns in Discovery Tooling”

During Discovery, tooling commonly applies the following analysis patterns:

  • Record normalization: Merge heterogeneous exports into one coherent application model.
  • Dependency mapping: Detect communication paths, data exchange, and coupling patterns.
  • Criticality and risk scoring: Evaluate business impact, failure domain, and compliance exposure.
  • Utilization profiling: Build workload demand baselines for right-sizing and target planning.
  • Segmentation analysis: Cluster applications by readiness, constraints, and migration strategy fit.
  • Wave simulation: Model move groups and sequence options under dependency constraints.
  • Gap and assumption tracking: Keep unresolved findings transparent with confidence levels.

These analyses establish the technical fact base. The human-driven stream then validates, prioritizes, and contextualizes these findings for executable migration decisions.

Use AI-assisted discovery assets to structure workload inputs, service mapping, readiness findings, and R-strategy signals before architects validate the resulting discovery baseline.

Asset title
Framework
Asset type

  1. Aggregate source data from CMDBs, hypervisors, cloud inventories, monitoring, and export files.
  2. Normalize and consolidate records into a common application-centric model.
  3. Discover and validate dependencies (network, data, identity, integration, and batch flows).
  4. Enrich with owner input on criticality, lifecycle, constraints, and migration feasibility.
  5. Classify workloads for migration strategy options and wave sequencing.
  6. Validate findings with architecture, security, platform, and business stakeholders.
  • Reduces migration risk: Early visibility of hidden dependencies lowers outage and rollback risk.
  • Improves wave planning: Workloads can be grouped realistically by coupling, criticality, and readiness.
  • Prevents over/under-sizing: Measured utilization replaces assumptions in target capacity planning.
  • Supports governance: Security, compliance, and operational constraints are addressed before rollout.
  • Strengthens stakeholder buy-in: Shared facts improve decision quality across business and IT.

Discovery outputs are directly reused by the next modules in Design and Mobilize:

Design

Uses dependency, capacity, and risk insights to shape target architecture options.

Security and Compliance

Uses data classification and control gaps to define prioritized security requirements.

Landing Zone

Uses platform and governance constraints to define foundational setup decisions.

Migration Plan

Uses move groups, criticality, and sequencing constraints for realistic wave planning.

Operating Model and Business Case

Uses ownership, process impact, and value/risk signals for staffing and investment priorities.

At minimum, Discovery should produce the following outputs:

  • Consolidated application baseline: Mapped inventory by domain, environment, and criticality.
  • Dependency map: Verified upstream/downstream relationships and integration touchpoints.
  • Utilization profile: Evidence-based resource behavior and sizing assumptions.
  • Constraint register: Security, compliance, licensing, and operational constraints.
  • Migration readiness view: Prioritized candidates, risks, and sequencing recommendations.

These outputs are essential prerequisites for continuing with detailed design work and a credible migration plan.

PLAN

Classify the Workload

Before choosing how to migrate, make transparent what is migrated. Use the classification dimensions to profile migration object, state, criticality, connectivity, downtime, and compliance per request.

Design and mobilizeDesignUse Cases In 2 trails

R-strategy explains how migration is run. This page adds the workload lens and clarifies what is migrated. It is intentionally solution-neutral and focuses on transparent categorization, feasibility boundaries, and decision context.

Application runtime migration

Migration of complete application runtimes across VM and platform targets, including dependencies, cutover behavior, and operating handover.

Container platform migration

Migration between Kubernetes platforms with separate treatment of stateless and stateful workload profiles.

Data migration

Migration of large file volumes, databases, and data platform workloads with explicit consistency, performance, and integrity boundaries.

Identity and access migration

Migration of IAM foundations such as SSO, federation, roles, service accounts, and permission models.

Network and connectivity migration

Migration of routing, DNS, firewall rules, segmentation, private connectivity, and cross-environment communication paths.

Integration and API migration

Migration of API contracts, messaging, eventing, and integration endpoints across source and target estates.

Security and compliance controls migration

Migration of controls, evidence chains, key material, and audit requirements required for regulated production readiness.

Operations and observability migration

Migration of monitoring, alerting, logging, incident workflows, and service-level operations baselines.

Delivery and resilience migration

Migration of CI/CD pipelines, automation controls, backup chains, and disaster recovery capabilities.

  • Application stack migration is part of Application runtime migration.
  • Large file data migration is part of Data migration.
  • Kubernetes to Kubernetes migration is part of Container platform migration.

Migration use cases from AWS and Azure to STACKIT

Section titled “Migration use cases from AWS and Azure to STACKIT”

STACKIT supports migration paths from AWS and Azure across application runtimes, data, network, identity, operations, and delivery. Select the target path for each workload according to its architecture, data profile, availability requirements, and operating model.

For the service-by-service target mapping, see AWS and Azure Target Service Mappings.

Use these dimensions to categorize requests before selecting implementation variants:

  • Migration object: Application stack, data set, or container platform.
  • State profile: Stateless, stateful, or mixed.
  • Criticality profile: Business criticality and accepted migration risk.
  • Connectivity profile: Network reachability and protocol compatibility between source and target.
  • Downtime profile: Allowed service interruption and cutover window constraints.
  • Compliance profile: Security, audit, and regulatory boundaries.
  1. Identify the primary migration object.
  2. Determine workload state profile and criticality.
  3. Capture connectivity and transfer constraints.
  4. Define boundary conditions for downtime, consistency, and compliance.
  5. Assign the request to a use-case category and record assumptions.
  • Includes: Application stack migration (including VM-centered migrations).
  • Primary objective: Transition application runtimes with predictable cutover and operating handover.

Mandatory boundaries:

  • Dependency clarity: Integration dependencies and transition windows must be known.
  • Cutover model: Interruption model and decision gates must be approved.
  • Rollback readiness: Triggers and ownership must be defined.
  • Includes: Kubernetes to Kubernetes migration.
  • Primary objective: Transition containerized workloads to target cluster models.

Mandatory boundaries:

  • State classification: Stateless/stateful boundaries must be explicit per component.
  • State portability: Storage and database compatibility must be validated.
  • Traffic control: Progressive switch capability must be confirmed.
  • Includes: Large file data migration and database/data platform transitions.
  • Primary objective: Transition data sets with controlled consistency and integrity.

Mandatory boundaries:

  • Connectivity feasibility: Required endpoint reachability must be validated.
  • Consistency model: Freeze windows, delta strategy, and validation methods must be defined.
  • Performance feasibility: Throughput profile and run limits must be validated.
  • Primary objective: Transition identity trust and access models without security regression.

Mandatory boundaries:

  • Trust model mapping: Federation, SSO, and token flows must be mapped.
  • Authorization mapping: Role and entitlement mapping must be validated.
  • Credential transition: Secret rotation and emergency access paths must be approved.
  • Primary objective: Transition communication paths and security boundaries between environments.

Mandatory boundaries:

  • Addressing and routing: IP planning and route ownership must be defined.
  • Control policy parity: Firewall and segmentation policies must be aligned.
  • Name resolution continuity: DNS transition behavior must be planned.
  • Primary objective: Transition service interfaces and integration patterns without breaking consumers.

Mandatory boundaries:

  • Contract compatibility: Versioning and compatibility strategy must be explicit.
  • Dependency sequencing: Producer/consumer switch order must be managed.
  • Message semantics: Ordering, retries, and repeat-safe behavior assumptions must be validated.

7. Security and compliance controls migration

Section titled “7. Security and compliance controls migration”
  • Primary objective: Preserve or improve control effectiveness and audit readiness during transition.

Mandatory boundaries:

  • Control mapping: Required controls and evidence points must be mapped.
  • Key and certificate handling: Cryptographic material transition must be governed.
  • Audit continuity: Logging and evidence retention obligations must remain intact.
  • Primary objective: Ensure operational control and incident response readiness after migration.

Mandatory boundaries:

  • Observability baseline: Metrics, logs, traces, and alerts must be active before cutover.
  • Operating ownership: On-call and escalation paths must be assigned.
  • Service objectives: SLO/SLA targets and thresholds must be defined.
  • Primary objective: Transition software delivery and resilience capabilities to target operations.

Mandatory boundaries:

  • Pipeline continuity: CI/CD and release controls must remain auditable.
  • Recovery readiness: Backup/restore and DR assumptions must be validated.
  • Automation safety: Guardrails for deployment automation must be in place.
  • Downtime target: Planned downtime window versus continuous availability expectation.
  • Consistency requirement: Eventual consistency, near-real-time, or strict transactional consistency.
  • Change tolerance: How much architecture and application change is acceptable in the current wave.
  • Automation level: Manual, semi-automated, or fully automated run path.
  • Risk and reversibility: Ability to detect failure fast and revert without business-critical impact.

The implementation details are maintained in dedicated asset pages. This keeps the overview page category-focused and allows concrete templates to evolve independently.

Asset title
Framework
Asset type

Large data migration to STACKIT file service

Section titled “Large data migration to STACKIT file service”
Asset title
Framework
Asset type

Asset title
Framework
Asset type

STEP

Decide the R-Strategy

With workload profiles in hand, route each application into its migration path. The decision criteria make explicit when Relocate, Rehost, Replatform, Repurchase, Refactor, or Retain/Retire is the right call.

Design and mobilizeDesignOverview In 2 trails

The Design module creates the executable migration design for each application identified in Discovery. It is not a generic architecture exercise. The target is a concrete design package that a migration factory can run with predictable quality.

Target design per application

Defines workload architecture, service choices, integration approach, and constraints in the STACKIT context.

R-strategy-backed decision record

Documents the selected migration strategy and why alternatives were rejected.

Factory-ready migration runbook

Provides a step-by-step procedure for run teams, including rollback and validation checkpoints.

Handover package

Delivers all required inputs to Migration Factory Setup, Landing Zone, and Migration Plan.

These modules are connected, but they have different responsibilities:

Design (this module)

Decides target design and migration strategy per application and creates executable runbooks.

Migration Factory Setup

Enables delivery by selecting and preparing the right factory model, partner setup, and tooling stack.

Landing Zone

Provides the platform foundation and governance controls that target designs must comply with.

Migration Plan

Converts completed designs into realistic waves, sequencing, dependencies, and delivery milestones.

The R-strategy model is the core decision framework in this module. For each application, the selected R-strategy must be justified with architecture, business, risk, and operability evidence.

Swipe sideways to see the whole diagram
R-strategy migration method Decision flow from discovery to production with the seven R-strategies: Relocate, Rehost, Replatform, Repurchase, Refactor, Retain, and Retire. R-strategy migration methodFrom discovery and path selection through the seven R-strategies to validation, transition, and production.DiscoveryDiscoveryAssess / prioritizeAssess / prioritizeDetermine migration pathDetermine migration pathValidationValidationTransitionTransitionProductionProductionRelocateRelocate(move VM)Define Landing ZoneDefine Landing ZoneUse migration toolsUse migration toolsAUTOMATEMANUALInstallInstallConfigConfigDeployDeployValidation & handoverRehostingRehosting(move application)Define Landing ZoneDefine Landing ZoneUse migration toolsUse migration toolsAUTOMATEMANUALInstallInstallConfigConfigDeployDeployReplatformingReplatforming(lift and reshape)Define Landing ZoneDefine Landing ZoneMap Target PlatformMap Target PlatformAdapt Platform StackAdapt Platform StackRepurchasingRepurchasing(replace, drop and shop)Purchase COTS/SaaS and licensingPurchase COTS/SaaS and licensingMigrate business processMigrate business processRefactoringRefactoring(re-architecting applications)Redesign application/ infrastructure architectureRedesign application/ infrastructure architectureApp code developmentApp code developmentFull ALM/SDLCFull ALM/SDLCIntegrationIntegrationRetain/moveRetain/movekeep for now or move laterRetire/decommissionRetire/decommissionLanding zone foundationLanding zone foundationShared platform base for all paths
  • Relocate: Choose when moving virtualization stacks is faster than redesign and governance constraints still allow conversion to cloud-native operations.
  • Rehost: Choose for low change tolerance and strict timelines, where speed is prioritized over immediate modernization.
  • Replatform: Choose when moderate changes unlock major benefits through STACKIT managed platform capabilities.
  • Repurchase: Choose only for services available in the STACKIT ecosystem, including SaaS options such as ServiceNow, SAP, and partner offerings when they better fit business needs.
  • Refactor: Choose for strategic applications where cloud-native redesign creates clear value in resilience, agility, or cost profile.
  • Retain: Choose retain when timing or dependencies block migration now.
  • Retire: Choose retire when business value no longer justifies operational effort.

The diagram is not only an orientation aid. It is the shared decision and handover spine across Design, Landing Zones, Migration Factory Setup, and Migration Plan.

Consolidate inventory, dependency map, risk profile, and non-functional constraints before strategy branching starts.

Prioritize workload candidates by criticality, effort, and wave feasibility to focus design capacity where delivery risk is highest.

Select the right R-strategy per workload and route into the dedicated strategy page with explicit rationale and alternatives considered.

Run a shared quality gate across architecture, security/compliance, runbook quality, and operating readiness before approving wave execution.

Prepare controlled handover into migration execution with release readiness, communication model, and clearly assigned ownership for wave run and escalation paths.

Finalize production handover criteria and Day-1 operating baseline so migrated workloads enter run operations with clear accountability and evidence.

How to build the target design per application

Section titled “How to build the target design per application”
  1. Confirm the application scope and baseline from Discovery (dependencies, usage, criticality, constraints).
  2. Define business and technical design goals, including availability, security, compliance, and performance targets.
  3. Evaluate the R-strategy options against objective criteria and record the selected strategy with rationale.
  4. Map the target to STACKIT products and platform capabilities, including networking, identity, data, and operations patterns.
  5. Specify migration approach details (cutover model, data movement, integration transition, rollback strategy).
  6. Create the migration run book for the factory with explicit tasks, quality checks, and acceptance criteria.
  7. Validate design assumptions with architecture, security, platform, and business owners.
  8. Hand over the approved design package to Migration Factory Setup and Migration Plan.

At minimum, each application package should include:

  • Target architecture definition: Workload placement, service mapping, integration model, and non-functional requirements.
  • R-strategy decision record: Selected strategy, decision criteria, alternatives considered, and key risks.
  • Migration run book: Sequenced run steps, checks before run, rollback path, validation steps, and go-live criteria.
  • Dependency and interface impact: Required coordination with upstream/downstream systems and transition windows.
  • Compliance and security controls: Mandatory controls and evidence requirements for release readiness.

Cloud design patterns for typical STACKIT applications

Section titled “Cloud design patterns for typical STACKIT applications”

Use a dedicated pattern page to define a concrete target architecture before selecting the migration runbook.

Workload use-case lens for migration variants

Section titled “Workload use-case lens for migration variants”

In addition to R-strategy, use the workload lens to classify what is migrated and to select feasible implementation variants with explicit boundary conditions.

AI-assisted design assets support architects in turning application requirements, source-service context, and R-strategy options into reviewable target-design proposals. Use their outputs as input for architecture, security, platform, and business validation before approving a migration path.

Asset title
Framework
Asset type

Asset title
Framework
Asset type

Migration planning quality depends on design quality. Wave sequencing, factory throughput, and delivery risk are directly influenced by how precise the target designs and run books are. In practice, incomplete designs lead to unstable waves and avoidable delivery delays.

BASE

Automate the Landing Zone Foundation

Architecture overview of the STACKIT Landing Zone Accelerator, from bootstrap through platform capabilities to application landing zones and workloads

This asset provides a reusable foundation for implementing a STACKIT platform landing zone. It is designed for enterprise environments that need a structured baseline for governance, security, networking, cost controls, and automation.

Repository:

This repository is a single root-module accelerator with modular submodules.

  • The root module in src/main.tf orchestrates all platform and landing-zone building blocks.
  • Configuration is provided through flavor-specific variable files in src/config/.
  • You can deploy with OpenTofu or Terraform (same module graph).
  • The initial deployment uses a temporary bootstrap service account, then migrates to a managed backend and managed credentials.

The repository provides eight reference configurations in src/config/. Choose the simplest topology that satisfies the required network, security, organizational, tenant, and regional boundaries.

Standalone topology with a management foundation, sandbox, and public application landing zone

Use standalone.tfvars for the smallest foundation: governance, management, a sandbox, and a public application landing zone with its own network and direct internet access. It creates no shared Network Area or connectivity hub, making it suitable when workloads do not require private east-west connectivity or central DNS.

Hub-and-spoke topology with a shared Network Area and separate public landing zone

Use hub-and-spoke.tfvars when corporate workloads need shared private connectivity. A central connectivity project provides the Network Area and DNS, the corporate data platform joins that private domain, and public workloads retain independent networks with direct internet access.

Hub-and-spoke topology with centralized OPNsense firewall inspection

Use hub-and-spoke-firewall.tfvars when corporate egress needs a consistent inspection and control point. It extends the shared Network Area with an OPNsense firewall and steers corporate default routes through the appliance, while public landing zones remain directly connected.

Finance and research topology with independent private connectivity domains

Use hub-and-spoke-finance-research.tfvars when business units require independent ownership and private connectivity. Finance and research each receive their own address plan, connectivity project, Network Area, and workload landing zone within the same STACKIT organization.

Multi-area topology separating regulated and shared workloads

Use hub-and-spoke-multi-area.tfvars when regulated and shared workloads must occupy separate private connectivity domains. Each domain has its own Network Area and DNS zone, with no implicit routing between them.

Multi-region topology with independent hubs in eu01 and eu02

Use hub-and-spoke-multi-region.tfvars for regional foundations in eu01 and eu02. Each region receives an independent hub, Network Area, workload landing zone, and optional platform Kubernetes cluster. Inter-region connectivity is deliberately not created and must be designed explicitly.

Production and non-production topology with separate Network Areas and firewalls

Use hub-and-spoke-prod-nonprod-firewall.tfvars when production must be isolated from non-production. Each domain receives a dedicated Network Area and OPNsense firewall; development and test share the non-production domain while remaining separate landing zones.

Tenant isolation topology with three independent private tenant domains

Use hub-and-spoke-tenant-isolation.tfvars for multiple tenants inside one organization. Every tenant receives an independent owner, address plan, Network Area, connectivity project, and workload landing zone, without private routing to the other tenant domains.

Code & registry github.com Landing Zone Accelerator architecture Review the implementation architecture, deployment configurations, and network behavior in the source repository. Open the repository

1. Governance module (src/modules/governance)

Section titled “1. Governance module (src/modules/governance)”

Purpose:

  • Creates RM folder structure (platform, landing_zones_corporate, landing_zones_public, sandboxes).
  • Assigns folder-level owners and auditors.
  • Assigns organization-level owners and auditors.
  • Creates custom roles at organization scope.

Landing-zone classification:

  • Platform Landing Zone: core governance baseline.

2. Management module (src/modules/management)

Section titled “2. Management module (src/modules/management)”

Purpose:

  • Creates a central management project.
  • Provisions Secrets Manager and default access user.
  • Provisions object storage buckets (including a remote state bucket).
  • Creates object-storage credentials and stores them in Secrets Manager.
  • Creates automation service account + rotating keys, stores key in Secrets Manager.
  • Optionally provisions observability and stores observability credentials in Secrets Manager.
  • Optionally configures federated identity providers for the automation service account.

Landing-zone classification:

  • Platform Landing Zone: shared operations and automation control plane.

3. Connectivity module (src/modules/connectivity)

Section titled “3. Connectivity module (src/modules/connectivity)”

Purpose:

  • Creates a dedicated connectivity project.
  • Creates network area and regional network-area configuration.
  • Creates DNS zones for shared naming domains.
  • Optionally creates firewall image, volume, server, interfaces, public IP.
  • Exposes firewall next-hop IP for route injection into corporate landing zones.

Landing-zone classification:

  • Platform Landing Zone: shared network and routing baseline.

Purpose:

  • Creates a dedicated DevOps project.
  • Optionally creates a central STACKIT Git instance with ACL ranges.

Landing-zone classification:

  • Platform Landing Zone in this accelerator’s architecture.
  • Rationale: it provides shared delivery tooling and central CI/CD source-control capability across landing zones.

5. Landing-Zone module (src/modules/landing-zone)

Section titled “5. Landing-Zone module (src/modules/landing-zone)”

Purpose:

  • Creates application-facing landing-zone projects (iterative via for_each).
  • Supports corporate landing zones (network area-connected) and public landing zones.
  • Creates routed networks, optional routing-table default route via firewall next hop.
  • Optionally creates per-project child DNS zones.
  • Creates project-level custom roles and role assignments.
  • Creates Secrets Manager, object-storage buckets, and automation service-account key material per landing zone.

Landing-zone classification:

  • Application Landing Zone: primary ALZ implementation module.

6. Sandboxes module (src/modules/sandboxes)

Section titled “6. Sandboxes module (src/modules/sandboxes)”

Purpose:

  • Creates lightweight sandbox projects in the dedicated sandboxes folder.
  • Assigns project owners.

Landing-zone classification:

  • Application Landing Zone (supporting): non-production experimentation space close to ALZ usage patterns.

The current implementation now includes an end-to-end path for a central Kubernetes platform and namespace-based application onboarding.

Platform landing zone for central Kubernetes

Section titled “Platform landing zone for central Kubernetes”

The platform scope now includes a dedicated central Kubernetes foundation that can be operated as a shared platform service.

  • Central cluster foundation: A dedicated platform Kubernetes project with SKE cluster life cycle, DNS extension integration, and optional observability wiring.
  • Secrets policy readiness: Namespace-level Secret Manager policy enforcement supports staged rollout modes such as audit and strict.
  • Shared service model: Platform teams can expose central capabilities while keeping project and namespace boundaries explicit.
  • Operational baseline: Cluster-level outputs and access information are exposed for automation and controlled platform operations.

Application landing zone for namespace tenants

Section titled “Application landing zone for namespace tenants”

Application landing zones can now consume namespace service from the central Kubernetes platform cluster.

  • Namespace onboarding: Landing-zone configuration can request namespace creation for an application team in the shared cluster.
  • Developer access path: Namespace-scoped Kubernetes users and role bindings are provided for tenant-level operations.
  • Service exposure: DNS and ingress patterns are preconfigured for application endpoints based on landing-zone and namespace context.
  • Secrets integration: Workloads can consume centrally governed secret flows while staying in namespace scope.

Extra platform features used in this setup

Section titled “Extra platform features used in this setup”
  • External DNS automation: DNS records for namespace services are managed from Kubernetes annotations and extension-zone integration.
  • Central Kubernetes monitoring: Platform observability integration includes Grafana access and metrics push wiring for cluster-level telemetry.
  • Dashboard provisioning workflow: Example dashboards are provisioned and imported for faster operational handover.
  • Encrypted volumes option: Platform module support for encrypted volume patterns helps align with stricter data-protection requirements.
  • Flexible network posture: The platform Kubernetes module supports SNA-oriented network setups for controlled enterprise connectivity.

What developers get in an application landing zone

Section titled “What developers get in an application landing zone”
  • Ready-to-use namespace: A preconfigured namespace in the central cluster instead of a full cluster-per-team model.
  • Least-privilege access: Namespace-scoped identities and permissions aligned to day-2 developer tasks.
  • Consistent endpoint model: Predictable DNS and ingress patterns for service publication.
  • Governed secret usage: Central secret governance with namespace-level consumption patterns.
  • Observability visibility: Shared metrics and dashboard views that help teams validate rollout and runtime behavior.

Platform vs application landing zone scope in this repository

Section titled “Platform vs application landing zone scope in this repository”
  • Platform Landing Zone focus (majority of implementation):
    • governance
    • management
    • connectivity
    • devops
  • Application Landing Zone scope (narrower by design):
    • landing-zone (core ALZ provisioning)
    • sandboxes (supporting ALZ-adjacent environments)

This means the repository delivers a strong platform baseline first, while ALZ capabilities are intentionally focused on project landing-zone provisioning and team sandbox enablement.

How to use this asset in migration programs

Section titled “How to use this asset in migration programs”
  • Start the platform baseline early (governance, management, connectivity, optional DevOps).
  • Define corporate vs public ALZ patterns based on connectivity and compliance needs.
  • Instantiate application landing zones with the landing-zone map in your variable files.
  • Use sandboxes for team onboarding and controlled early experiments.
  • Move state to the managed backend after first apply and switch from bootstrap credentials to managed automation credentials.
  • Reusable baseline modules: Building blocks for account/project structure, IAM, network, and controls.
  • Policy-oriented setup: Guardrails and conventions for secure and governed cloud usage.
  • IaC-first approach: OpenTofu/Terraform implementation model for repeatable provisioning.
  • Enterprise extensibility: Designed as a baseline to be adapted to customer-specific requirements.
  • Early platform stream: Start foundation setup in parallel with discovery.
  • Control baseline before production move: Ensure mandatory controls are in place before productive migrations.
  • Template source for application landing zones: Reuse and refine baseline modules for workload archetypes.
  • Organization and ownership model for projects/environments.
  • Security and compliance requirements (identity, logging, evidence, segmentation).
  • Connectivity constraints and integration requirements.
  • Operating model alignment across platform, security, and application teams.
SAFE

Anchor Security and Compliance Early

Security and compliance is a parallel stream, not a final gate. Map the control topics early so evidence pipelines, sovereignty requirements, and zero-trust decisions land in the design instead of the audit.

Security and Compliance Choose an entry point by topic cluster and jump directly to the detailed module page. Security and ComplianceChoose an entry point by topic cluster and jump directly to the detailed module page.Governance and migration operating modelOperating model and governanceOperating model and governanceRoles, checkpoints, ownership, and decisionsOn-premises to cloud shiftOn-premises to cloud shiftControl translation and responsibility shiftsSecurity architecture and design baselineArchitecture patternsArchitecture patternsBoundary-centric and Zero Trust combinationsSecurity by design baselineSecurity by design baselineMandatory baseline controls by domainCompliance assurance and sovereigntyControls and evidence pipelineControls and evidence pipelinePreventive, detective controls and evidenceDigital sovereignty and CSFDigital sovereignty and CSFCSF alignment, ES3, and auditabilityZero Trust deep dive topicsFive focused domains for implementation and controls.Zero trust peopleZero trust peopleIdentity lifecycle, privileged access, and account hygieneZero trust devicesZero trust devicesDevice posture, endpoint hardening, and secure administrationZero trust networksZero trust networksSegmentation, traffic policy, and controlled connectivityZero trust workloadZero trust workloadRuntime hardening, least privilege, and workload isolationZero trust dataZero trust dataClassification, encryption, keys, and retention controls
Security and compliance topic map
LIFT

Set Up the Migration Factory

Scale from single moves to governed waves. The factory operating model defines intake, roles, throughput, and quality gates so migration capacity grows across teams without losing control.

Design and mobilizeMigration planOverview In 2 trails

Migration Plan is the transition from analysis to delivery. It follows Discovery and is continuously filled with concrete implementation detail from the Design module.

The goal is to transform findings from Discovery and Design into a detailed, step-by-step migration plan that delivery teams can run with predictable quality.

This module forms the bridge between strategy and implementation and is therefore critical for the overall migration success.

Wave planning

Group applications into logical migration waves based on dependency constraints, business criticality, and technical complexity.

Migration runbooks

Create detailed, step-by-step runbooks per wave or application covering preparation, delivery, cutover, rollback, and post-migration validation.

Resource planning

Define required teams, skills, tools, and expert allocation per wave, including enablement and training planning.

Delivery governance

Establish clear governance processes, ownership, communication cadence, and cutover controls for wave delivery.

Migration planning must stay adaptive. Programs should start initial waves as early as possible and then continuously refine wave scope and sequencing as additional Discovery and Design outputs become available.

Runbooks are living documents. Teams should review and improve runbooks after every cutover. As runbook quality increases, migration velocity typically increases wave by wave.

In large migration programs, the goal is to scale delivery from initial pilot waves to predictable factory throughput. A typical trajectory is to start with small waves (for example, 5 servers/week) and gradually increase throughput (for example, up to 50-100 servers/week), depending on constraints and delivery maturity.

Early waves are intentionally smaller so portfolio and migration workstreams can stabilize their processes, validate assumptions, and improve runbooks. This learning loop is a key success factor for large migrations.

In this module, the migration factory is typically operated through four components:

Project governance rules

Processes and tools that govern wave orchestration, communication, timelines, and cutovers so teams run tasks in the right sequence and at the right time.

Portfolio runbooks

Runbooks used to prioritize applications, plan waves, and collect migration metadata as the input material for delivery.

Migration runbooks

Runbooks used to run migration waves, load metadata into migration tooling, and complete cutover and validation.

Best practices and health-check matrix

A regular health-check mechanism used to assess progress, identify delivery risks early, and keep delivery on track.

Data Flow Through Portfolio and Migration Workstreams

Section titled “Data Flow Through Portfolio and Migration Workstreams”

Runbooks create the data flow across two connected workstreams:

  • Portfolio workstream: Prioritizes and prepares applications and metadata for upcoming waves.
  • Migration workstream: Runs migrations and cutovers according to approved wave plans.

Teams are usually dedicated to parts of the factory while waves flow through both workstreams. To prevent supply issues, keep enough prepared waves in front of delivery. A common baseline is to keep the portfolio workstream five waves ahead of the migration workstream.

For governance and capacity planning, it is important to separate function from team ownership:

  • Portfolio: Functional and technical wave preparation, including prioritization, dependency clarification, scope shaping, metadata quality, and readiness proof.
  • Portfolio Team: Roles that run and own portfolio work (for example, program leadership, domain owners, architects, application owners, and governance).
  • Migration: Operational delivery of approved waves, including runbook-driven delivery, change/cutover control, validation, stabilization, and documented closure.
  • Migration Team: Roles that run and secure technical migration delivery (for example, factory engineers, platform teams, network/security specialists, test, and operations handover).

Both teams are tightly coupled, but they deliver different outputs:

  • Portfolio Team delivers: Approved wave cuts, prioritized backlogs, complete migration metadata, and delivery-ready input packages.
  • Migration Team delivers: Successful cutovers, validated target states, lessons learned, and improved runbook versions for upcoming waves.

A typical dynamic pattern is:

  • Portfolio cadence: About 1-2 weeks per wave for preparation.
  • Migration cadence: About 3-4 weeks per wave for delivery and cutover.
  • Wave buffer: A five-wave buffer between portfolio and migration workstreams.

Before migration delivery reaches steady throughput, portfolio planning usually establishes an initial wave buffer. Once delivery starts, both workstreams continue in parallel and the buffer helps prevent delivery stalls.

The following diagram visualizes the operating model based on phases and modules: planning remains anchored in the Migration Plan module (Design and Mobilize), while delivery runs in staggered migration waves.

In this model, phase boundaries are intentionally overlapping:

  • Design and Mobilize starts with wave planning and remains active through the planning of Wave 8 (up to Week 3).
  • Migrate starts already in Week 3 and continues through the remaining timeline.
  • Pilot wave and adjustments: Wave 1 is a shorter pilot wave for initial validation, followed by targeted setup adjustments during the first production waves.

The detailed setup scope is documented in its dedicated chapter: Migration Factory Setup.

Swipe sideways to see the whole diagram
Wave model across phases and modules Timeline of wave-based planning and migration execution with Portfolio Team and Migration Team handover pattern. Design and Mobilize phase Migrate phase Migration Factory Setup Factory Setup Adjustment Small Factory Adjustments Week 1 Week 2 Week 3 Week 4 Week 5 Week 6 Week 7 Week 8 Week 9 Wave 1 Wave 2 Wave 3 Wave 4 Wave 5 Wave 6 Wave 7 Wave 8 PT Plan PT Plan PT Plan PT Plan PT Plan PT Plan PT Plan PT Plan MT Pilot wave MT Migration MT Migration MT Migration MT Migration MT Migration MT Migration MT Migration PT Portfolio Team MT Migration Team

The pattern is intentionally dynamic: portfolio work keeps a stable forward buffer, while migration work runs waves with runbooks, cutovers, and validation.

Migration Factory Process in the Module Flow

Section titled “Migration Factory Process in the Module Flow”

Migration Plan connects Discovery and Design outputs with delivery governance, wave planning, and runbook-driven delivery. The process tightly links portfolio and migration workstreams through governance controls, runbooks, and continuous improvement.

Wave planning is not a static schedule. It is an operational timeline that must remain transparent for stakeholders and adaptable to new findings. Cadence, overlap between workstreams, and buffer logic are the key control variables.

  1. Confirm latest Discovery and Design inputs, constraints, and assumptions.
  2. Update wave composition, sequencing, and staffing based on current facts.
  3. Run the current wave with approved migration and cutover runbooks.
  4. Review cutover outcomes, incidents, and timing variances.
  5. Improve governance controls and runbooks, then apply updates to upcoming waves.
  6. Re-check health status and wave buffer, then continue with the next iteration.

At minimum, Migration Plan should produce:

  • Approved wave plan: Sequenced migration waves with dependency-aware grouping.
  • Executable runbooks: Versioned runbooks per wave/application with validation and rollback.
  • Resource and skill plan: Staffing and capability mapping per wave.
  • Governance cadence: Decision, communication, and cutover control framework.
  • Continuous improvement backlog: Tracked runbook and process improvements from each wave.
AUTO

Run the Migrate Phase

Execute approved waves through the core modules of the Migrate phase, with clear boundaries between migration runs, optimization, and the operating-model handover.

Overview In 1 trail

Migrate is the phase where migration happens for R-strategy paths that involve a technical move. It uses a factory-based delivery model to move workloads in controlled waves from source environments to STACKIT target environments.

The phase is runbook-driven and focuses on repeatability, cutover quality, and stable handover into operations.

Migrate does not look the same for every R-strategy decision. The phase focus depends on the selected path per workload:

  • Relocate, Rehost, Replatform: These are the core migration paths in this phase and use wave delivery with runbooks.
  • Refactor: This is usually run as a dedicated modernization project stream with its own backlog and delivery setup.
  • Repurchase: This follows a different transition pattern and is often best handled as a dedicated workstream or module.
  • Retain, Retire: These paths do not require migration run in this phase and are treated as portfolio decisions.

Migrate starts when the first migration waves are approved and their runbooks are ready.

  1. Start with validated wave scope, target architecture decisions, and operational cutover plans.
  2. Execute migration waves with governance controls, technical validation, and rollback readiness.
  3. End per workload group when migrations are completed and post-cutover stabilization is accepted.

At program level, the phase ends when all planned workloads are migrated, modernized, or formally retained/retired according to scope decisions.

Migrate is where migration value is realized in production. Delivery quality in this phase directly impacts business continuity, user trust, and long-term operating efficiency.

Controlled cutovers

Standardized runbooks and wave governance reduce outage and rollback risks.

Factory scale

Reusable migration patterns increase throughput while preserving quality.

Measured validation

Technical and functional checks confirm target-state stability after each wave.

Continuous improvement

Lessons learned are fed back into upcoming waves, optimize cycles, and modernization paths.

  • Migrate: Delivers workload relocation and cutover in controlled migration waves.
  • Optimize: Improves workload sizing, performance, and cost after migration.
  • Refactor: Enables deeper restructuring and cloud-native improvements where required.
  • Repurchase: Covers SaaS replacement transition patterns that differ from technical migration waves.
  • Operating Model Handover: Validates Target Operating Model ownership and transfer readiness before full Run operations.

At the end of Migrate, the program should have:

  • Successful wave run: Approved waves completed with documented cutover and validation evidence.
  • Stabilized target workloads: Migrated systems running reliably on STACKIT with defined ownership.
  • Optimization backlog and actions: Rightsizing and performance improvements prioritized and implemented.
  • Modernization decisions: Refactor paths identified and initiated where business value is highest.

Migrate includes immediate post-cutover stabilization and optimization loops. Long-term service operations, life cycle governance, and sustained value realization are anchored in the subsequent Run phase.

LIVE

Cut Over with a Governed Runbook

  • Category: Application stack migration
  • Typical source: Existing VM-based runtime
  • Typical target: VM runtime on STACKIT with standardized operations handover
  • Dependency map completed: Interfaces, schedules, and critical upstream/downstream systems are documented.
  • Landing zone readiness: Network, IAM, monitoring, backup, and logging controls are available.
  • Cutover governance: Approved window, freeze rules, and business sign-off path are defined.
  • Rollback readiness: Source restore path and technical rollback trigger are tested.
  • Heavy redesign is required: Significant architecture changes are expected in the same wave.
  • No rollback option exists: Source cannot be kept stable for fallback during cutover.
  1. Provision target VM baseline and apply security controls.
  2. Install runtime dependencies and deploy the application artifact.
  3. Configure environment, certificates, and endpoint integration.
  4. Validate observability baseline (metrics, logs, alerts).
  1. Freeze non-essential writes on the source.
  2. Run final data synchronization.
  3. Validate integrity and application readiness on target.
  1. Stop source service according to cutover governance.
  2. Activate target service and verify health checks.
  3. Activate the approved target endpoint or traffic path and run business-critical validation.
  4. Start stabilization watch and complete evidence log.
  • Functional readiness: Core user journeys pass.
  • Data readiness: Consistency checks meet acceptance criteria.
  • Security readiness: Access controls and TLS chain verified.
  • Operations readiness: Runbook, alerts, and escalation ownership confirmed.
OPS

Stabilize Through Hypercare

Directly after cutover, run a focused Hypercare window: heightened monitoring, fast remediation, and explicit exit criteria close the immediate post-migration risks before steady-state operations.

RunHypercareOverview

Hypercare is the controlled stabilization period directly after migration cutover. It keeps migration and project teams close to operations so unresolved defects can be fixed early with low risk.

This module forms the explicit bridge from Migrate to Run.

  • Fast remediation window: Deep workload context is still available in delivery teams.
  • Risk containment: Early incidents can be resolved before they become recurring operational debt.
  • Operational hardening: Missing monitoring, alerting, and runbook details are completed under real load.
  1. Confirm cutover baseline and define Hypercare entry criteria.
  2. Track incidents, instability patterns, and operational blind spots daily.
  3. Prioritize remediation by business impact and service criticality.
  4. Implement fixes with controlled change windows and rollback readiness.
  5. Verify stabilization criteria and close unresolved risks with clear ownership.
  6. Exit Hypercare when service, observability, and support readiness are accepted.
  • Post-cutover defect closure: Resolve migration-related errors and service degradations.
  • Monitoring completion: Add missing metrics, dashboards, and alert thresholds.
  • Runbook maturity: Refine troubleshooting, escalation, and rollback procedures.
  • Knowledge transfer preparation: Structure findings for operating model handover.

Primary inputs

Cutover records, migration runbooks, open issue logs, SLO baselines, and incident patterns.

Hypercare outputs

Stabilized services, closed critical defects, completed observability baseline, and handover-ready documentation.

Exit result

Approved transition into Operate and regular run governance.

GOAL

Hand Over to Customer Success

Close the expedition by connecting the stabilized workload to the long-term coordination layer: adoption tracking, stakeholder alignment, and escalation routing through STACKIT Customer Success.

RunCustomer SuccessOverview

Customer Success is an optional module for programs that want a dedicated coordination role for adoption and outcome realization in Run.

A Customer Success Manager can act as central contact point between customer stakeholders, platform teams, support channels, and operating governance.

At STACKIT, Customer Success Managers proactively support customers in maximizing value and outcomes to ensure long-term satisfaction and retention.

Core responsibilities include guiding customers across the full customer journey, analyzing customer needs, solving issues proactively, and identifying growth potential to develop the business relationship strategically.

Customer Success serves as the customer contact on a meta level and as interface to all other STACKIT domains.

Through regular feedback loops and proactive communication, Customer Success ensures satisfaction while also acting as an escalation instance toward top-level management for strategic alignment.

  • Adoption orchestration: Align enablement activities, usage growth, and service onboarding.
  • Outcome tracking: Connect technical run metrics to business success indicators.
  • Cross-team coordination: Reduce friction between support, operations, platform, and customer stakeholders.
  • Improvement prioritization: Convert recurring issues and value opportunities into clear backlog priorities.
  1. Define success outcomes and stakeholder map for the run period.
  2. Establish regular cadences for service health, adoption, and value tracking.
  3. Coordinate escalations and follow-up actions across customer and provider teams.
  4. Review delivered outcomes and adjust priorities for the next operating cycle.

Outcome dashboards

Shared visibility across reliability, adoption, and business-relevant service KPIs.

Coordinated action backlog

Prioritized follow-up actions across support, operations, and service improvement.

Stakeholder alignment

Clear communication lines and decision cadence for run-related evolution topics.

Trail historyActive 4 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
Maintainers
TMTobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172Contributed in TM Solutions Corp. Inc.
Show full history (8 more)