Sovereign, secure cloud operations need end-to-end automation and regulatory resilience: from standardized IaC to automated compliance to disaster recovery.
PLAN
Infrastructure as Code: Strategic Maturity
Establish IaC as a strategic leadership decision and build IaC maturity step by step, from manual work through centralized state to automated compliance checks.
Adapting Operating ModelsInfrastructure-as-Code In 1 trail
What Infrastructure-as-Code means — and why it is a leadership decision
Infrastructure-as-Code is the practice of describing infrastructure — servers, networks, databases, security rules — in writing, rather than configuring it through manual clicks in a console. What sounds technical is, in its effect, a leadership decision: the organisation decides that infrastructure changes must be documented, traceable, and reviewable — like any other significant decision.
This decision has profound organisational consequences. Infrastructure documentation no longer becomes outdated, because the system always matches what has been written. New environments — a second location, an additional test environment, disaster recovery — are not the product of days of configuration work, but of repeating existing code. Errors are caught before they occur, because changes are made explicit and can be reviewed.
The decisions that come before the first line of code
Before the first infrastructure description file is written, three organisational questions must be answered. They determine whether IaC becomes a stable operating model or a well-intentioned approach that isn’t followed in practice.
Who writes and maintains the infrastructure descriptions? The two extreme answers are both problematic. If only the Platform Team writes code, a bottleneck emerges: teams must wait for infrastructure instead of working independently. If every team acts fully autonomously without common standards, an incompatible tangle of approaches emerges. The methodically correct path lies in the middle: the Platform Team sets standards and provides reusable building blocks. Teams work autonomously within those boundaries.
How are changes to production environments approved? An infrastructure change in production made by one person, without review, without a record — that is the same risk as clicking manually in the console, only harder to undo. The sensible requirement is that production environment changes always run through an approved automated process, never directly by hand.
Where is the system state stored? IaC tools manage state: what is currently deployed? This state is critical to the functioning of the entire approach. It must be stored securely, centrally, and accessibly for the team — never on one person’s laptop, never in a publicly accessible system.
Infrastructure-as-Code changes the requirements for people who manage infrastructure. This is not primarily a technical challenge — it is a learning curve that requires time, practice, and a fault-tolerant learning environment.
What changes
Systems administrators who have managed infrastructure through configuration interfaces for
years must develop a different mental model: instead of “I click Create,” they think “I describe
the desired state.” This shift is learnable, but it takes time and must not be underestimated.
What stays
The infrastructure knowledge stays — and is the biggest advantage. Someone who understands how
networks, security groups, and databases work learns the new description form quickly. Technical
knowledge doesn’t need to be rebuilt, only the method of applying it.
What a good learning environment needs
Sandbox access for experimentation without consequences, pair work with more experienced
colleagues, no punishment for mistakes in test environments. Someone who makes mistakes while
learning, learns. Someone who makes mistakes and is criticised for them stops learning.
The Platform Team's role
The Platform Team is not a control body in this phase — it is an enabler. It provides learning
materials, answers questions, reviews code alongside teams — not to evaluate, but to teach.
The transition to Infrastructure-as-Code is not a project with a start date and a sign-off. It is a maturity curve that unfolds over months and years. What matters is knowing the current position and taking the next sensible step — not trying to immediately reach the highest maturity level.
The first stage is the simplest: new resources are described as code, existing resources remain manual for now. This is incomplete, but a genuine start. Every new piece of infrastructure that emerges as code is a resource the team fully understands, has documented, and can restore.
The second stage brings all production resources under code control and introduces a central state store. Manual changes to production environments become the exception that must be documented.
The third stage connects IaC with automated quality and compliance checks. Changes are not only reviewed, but automatically evaluated against security and governance requirements before being executed.
Infrastructure-as-Code is not the goal — it is the means. The goal is a cloud environment that is traceable, secure, reproducible, and auditable. IaC is the most efficient path to get there.
Organisations that introduce IaC typically report three effects after six to twelve months: new environments emerge in hours, not days; infrastructure incidents are resolved faster because the path back to the baseline state is known; and audit requests are handled more efficiently because the documentation of state is created automatically.
These effects are more motivating than any abstract reference to best practices.
Define the starting point: Which part of the infrastructure to begin with? New resources in a non-critical environment are the lowest-risk entry point. Existing complex production infrastructure is not a good starting point.
Provide standards and building blocks: The Platform Team develops reusable descriptions for common infrastructure components. These building blocks are the lever that allows teams to start quickly without having to reinvent every step.
Accompany the upskilling: Teams receive learning time, sandbox access, and support. This step is frequently underestimated — and is the most common reason why IaC introductions stall.
Complete and evaluate the pilot: What worked? What was harder than expected? What needs to be adjusted? This evaluation is the foundation for the broader rollout.
Expand incrementally: Based on pilot experience, bring in further teams until all production resources are under code control.
BASE
CI/CD Pipelines & Quality Gates
Build automated deployment pipelines with integrated, automatic quality and compliance checks before going live.
Adapting Operating ModelsCI/CD Pipelines In 1 trail
Continuous Integration and Continuous Deployment are not tools. They are organisational principles — a decision about how an organisation develops, validates, and operates software.
The core idea is simple: instead of deploying changes to production infrequently, manually, and riskily, they are delivered frequently, automatically, and in a controlled manner. Every change goes through the same quality checks. Every change is documented. Every change can be reversed.
Organisations that take this path don’t deliver faster because they work less carefully. They deliver faster because they have automated the care.
Before a CI/CD pipeline is built, two cultural prerequisites must be in place. Without them, any technical implementation will fail against the organisation.
Tests are not optional additional work. A pipeline that checks automatically can only be as good as the tests it runs. If tests are seen as an annoying obligation to get through as quickly as possible, the pipeline becomes a formality — a green light that means nothing. Tests must be understood as an integral part of development work, not as overhead at the end.
Small, frequent changes rather than large, infrequent releases. The greatest risks in software changes come from the size of the change: the larger the package, the harder the debugging, the greater the impact when something goes wrong. CI/CD works best when teams learn to decompose changes into smaller, deliverable units. This is a way of working that must be learned — it is not natural for teams accustomed to monthly releases.
Quality gates in a pipeline are decisions, not technical defaults. What must pass before a
change reaches production? Which tests are mandatory? Which security checks are non-negotiable?
Setting these gates consciously — and documenting why they were set — is more important than the
technical configuration.
Who is responsible for pipeline operations?
A pipeline is itself software — it must be maintained, updated, and analysed quickly when
problems occur. If ownership is unclear, problems are ignored or bypassed. The Platform Team
should own the pipeline infrastructure; individual teams own the tests and checks that their
changes pass through.
How are failing pipelines handled?
How the organisation deals with red pipelines is a cultural question. Are failures fixed
immediately? Or do they accumulate because “the test has been failing for weeks but it was never
a real problem”? A failing pipeline that no longer alarms anyone has stopped providing safety.
What happens when something goes wrong?
A rollback process must be defined before it is needed — not in the moment when a critical
problem has occurred. How long does a rollback take? Who can trigger one? Which teams must be
informed? Answering these questions calmly in advance is much better than answering them under
pressure.
The greatest danger at deployment is irreversibility: a change that causes a problem, and no fast path back. Deployment strategies address this through controlled introduction of changes.
The simplest strategy — rolling changes out progressively to a growing proportion of users — makes it possible to detect problems before all users are affected. If something goes wrong, five percent of users are affected, not a hundred. Rollback is a decision, not a catastrophe.
A more sophisticated strategy — running a new version in parallel with the old one until the new version has proven it works — eliminates the risk of outage entirely. If the new version shows problems, the switch is made back to the old one. No user experienced an outage.
Which strategy is right for which context depends on the criticality of the workload and the maturity of the team. A highly critical production environment needs different safety mechanisms than an internal tool.
An often-overlooked value of CI/CD is documentation. Every change that runs through a pipeline leaves evidence: what was changed, when, by whom, what checks were performed, was the result positive?
For organisations under regulatory requirements, this is not a side effect — it is a compliance asset. Instead of laboriously assembling audit requests from logs and memory, there is a complete, immutable trail for every production change. This substantially reduces the effort of audits.
Choose a single pilot: Identify a workload where the team is motivated and willing to take risks. Not a core production service, but an important, non-critical one. Build the first pipeline together with the team — so the team understands and owns the pipeline, not just uses it.
Define quality gates: What should this pipeline check? What does it block? These decisions are discussed with the team and documented. No automatic adoption of defaults — every gate has a reason.
Evaluate the pilot phase: What worked well? What created friction? What should have been prevented? This review is the foundation for the broader rollout and should be conducted honestly.
Rollout with support, not pressure: Bring in more teams — with support, not mandate. Teams that experience pipelines as a constraint build workarounds. Teams that experience pipelines as protection maintain them.
Measure quality and learn: How long does a change take from development to production? How often does the pipeline fail and why? These metrics — not as a control mechanism, but as a learning source — show where investment should go.
CI/CD and Infrastructure-as-Code are not separate topics. In a mature environment, infrastructure changes are treated exactly like application changes: they run through a pipeline, are checked automatically, and changes to production environments happen exclusively through this approved process. This interplay — application code and infrastructure code in the same quality processes — is the hallmark of a mature cloud operating culture.
STEP
Landing Zone Governance
Provide standardized platform and application landing zones to inherit security requirements down to product teams.
A landing zone is the structured cloud foundation that defines how your organization operates on
STACKIT from day 1. It combines governance, identity, security, network design, cost controls,
and automation into one coherent baseline.
Without this foundation, migration waves typically stall due to missing approvals, inconsistent
controls, and repeated platform decisions.
The following visual summarizes the core components that should be addressed for a reliable
platform baseline. These six building blocks form a secure platform foundation:
Account Governance — project structure, hierarchy, and ownership model
Identity & Access Management — users, roles, and federated identity
Security & Compliance — guardrails, policies, and compliance controls
Network Architecture — segmentation, connectivity, and traffic control
Cost Management and Control — tagging, budgets, and spend visibility
Automation (IaC) — infrastructure as code, pipelines, and drift detection
Many organisations decide to outsource parts of their cloud operations to a managed service provider (MSP). The decision is often correct — operations become more reliable, and internal teams can focus on strategic topics.
What is frequently underestimated: governing the MSP relationship is itself a demanding task. Anyone who does not build a structured steering process gradually loses control of their own cloud environment — without noticing.
An MSP contract must cover more than price and scope. Critical contract elements:
Service Level Agreement (SLA):
Availability SLA for the managed service (typical: 99.5% to 99.9%)
Response time by severity (P1: 15 minutes, P2: 1 hour, P3: 4 hours)
Resolution time targets with escalation path
SLA reporting frequency and format
Penalties and credits:
Without economic consequences for SLA violations, SLAs have little steering effect. Typical: credits of 10–30% of the monthly service fee for demonstrated SLA shortfall.
Subcontractor provisions:
Which tasks may the MSP further delegate? Every sub-delegation must be transparent and must meet the same data protection and security standards — relevant for GDPR Art. 28 para. 4.
Exit provisions:
How is the handover to another provider or back to the internal team governed? Timelines, knowledge transfer obligations, documentation handover, access return.
MSP governance requires structure. Without regular review, the relationship drifts in a direction that the client does not notice — until there is a problem.
Swipe sideways to see the whole table
Review format
Frequency
Participants
Agenda
Operational status call
Weekly
IT lead + MSP delivery
Open incidents, tickets, ongoing changes
SLA review
Monthly
CIO + MSP management
SLA report, deviations, improvement measures
Strategic review
Quarterly
Board + MSP leadership
Roadmap, contract adjustments, make-or-buy
Security audit
Annual
CISO + MSP security
Penetration test results, certifications, incident report
Does the MSP have demonstrable experience with STACKIT? Are there reference customers from
similar industries? STACKIT-certified partners provide a structured entry point.
Compliance expertise
Does the MSP understand the regulatory requirements of your sector? GDPR, TISAX, BAIT, KRITIS —
an MSP without compliance expertise is unsuitable for regulated environments.
Transparent processes
Can the MSP present their incident response process, change management procedures and security
concept? Providers who evade these questions are a risk.
Exit readiness
A reputable MSP actively shapes the exit clause — because it knows that a good relationship is
the best client retention. Anyone who blocks exit provisions is creating dependency.
Define scope — What is outsourced, what remains internal? In writing, as the basis for the tender.
MSP tender or selection — At least three offers, a structured evaluation matrix, reference conversations with existing clients.
Contract negotiation — SLAs, penalties, subcontractor provisions, exit clause, GDPR data processing agreement. Legal counsel specialised in cloud contracts is recommended.
Implement access concept — Set up MSP accounts, define permission scope, activate audit logging, establish review process.
Establish review rhythm — Weekly, monthly and quarterly reviews anchored in the calendar. Create agenda templates.
Agree documentation requirements in writing — What does the MSP deliver, in what format, at what interval? As a contractual component, not a verbal assurance.
The SLA requirements for MSPs are derived from the DR tier classifications of the affected workloads. The internal Service Catalogue defines which services IT delivers itself and which the MSP assumes.
GOAL
Disaster Recovery and Business Continuity
Proactively define, document, and regularly simulate RTO and RPO targets for business-critical systems on STACKIT.
Adapting Operating ModelsDisaster Recovery In 1 trail
Disaster recovery is the discipline nobody practises — until they have to. Then every decision that could have been made earlier costs money and trust.
For most organisations, cloud migration is the moment when DR requirements are defined in writing for the first time. This is the opportunity: clearly, proportionately, and tested.
Every system has an implicit recovery time and an acceptable data loss threshold. Most organisations do not know these figures — and pay the price when an incident occurs.
Recovery Time Objective (RTO): How long may a system be unavailable after an outage? An RTO of 4 hours means: after at most 4 hours, the system must be functioning again.
Recovery Point Objective (RPO): How much data loss is acceptable? An RPO of 1 hour means: at most the last 60 minutes of data may be lost.
The important rule: The more aggressive the RTO and RPO, the higher the ongoing costs. An RTO of 15 minutes with an RPO of 5 minutes requires hot-standby environments. An RTO of 24 hours with an RPO of 12 hours can be met with daily backups and manual processes.
The right answer is: proportionate, not maximalist.
For regulated industries, DR is not a voluntary best practice but a regulatory requirement.
BAIT (banking): Requires documented contingency plans, regular tests, and failure scenarios for all material IT systems. Recovery targets must be approved by the management board.
KRITIS (critical infrastructure): Obliges KRITIS operators to business continuity management in accordance with BSI IT-Grundschutz or ISO 22301. DR tests are mandatory; results must be documented.
GDPR: Art. 32 requires the ability to restore the availability of personal data after an incident. An RPO of more than 24 hours for GDPR-relevant systems is difficult to justify.
DORA (Digital Operational Resilience Act, from January 2025): Obliges financial institutions to comprehensive ICT risk management with explicit BCM requirements and mandatory resilience tests.
STACKIT as a partner: STACKIT operates data centres in Germany with ISO 27001 and SOC 2 certification. Data residency in de-01/de-02 ensures that backup data does not leave Germany — a decisive advantage for regulated industries.
DR tests — why they rarely happen and how to change that
The most common DR mistake is never testing the plan. DR tests are regularly postponed because the production environment is considered too critical to risk a simulated failover, no clear ownership for DR tests exists, and tests are seen as effort without direct benefit.
The solution: Anchor DR tests as a regular engineering ritual, not as an exceptional event.
Four test formats in ascending intensity:
1. Tabletop exercise (annually, low effort): The team works through a failure scenario together — on the whiteboard. No systems touched. Identifies gaps in processes and accountabilities.
2. Component test (quarterly): A specific backup or failover mechanism is tested in the staging environment. Duration: a few hours.
3. Partial failover test (semi-annually): A non-production-critical system is actually switched to the standby site. Measures actual RTO. Duration: half a day.
4. Full DR exercise (annually): Complete simulation of a data centre outage for all Tier 1 and Tier 2 systems. Requires planning and alignment with stakeholders. Delivers the real evidence for regulators.
Approves RTO/RPO targets, acts as escalation authority in an emergency, signs off DR tests and
reports results to the board and regulators.
Platform Team
Implements and maintains DR infrastructure, conducts component tests, documents recovery
procedures.
Application owners
Classify their systems into tier classes, define application-specific recovery steps,
participate in tabletop exercises.
STACKIT
Guarantees availability of cloud infrastructure in accordance with SLA, provides zone-to-zone
replication, delivers audit logs for regulatory evidence.