Skip to content
Beta

Sovereign Engineering & Resilient Operations

Last updated on

Stackit LogoStackit Logo
STACKIT

Sovereign Engineering & Resilient Operations

Sovereign, secure cloud operations need end-to-end automation and regulatory resilience: from standardized IaC to automated compliance to disaster recovery.

PLAN

Infrastructure as Code: Strategic Maturity

Establish IaC as a strategic leadership decision and build IaC maturity step by step, from manual work through centralized state to automated compliance checks.

Adapting Operating ModelsInfrastructure-as-Code In 1 trail

What Infrastructure-as-Code means — and why it is a leadership decision

Section titled “What Infrastructure-as-Code means — and why it is a leadership decision”

Infrastructure-as-Code is the practice of describing infrastructure — servers, networks, databases, security rules — in writing, rather than configuring it through manual clicks in a console. What sounds technical is, in its effect, a leadership decision: the organisation decides that infrastructure changes must be documented, traceable, and reviewable — like any other significant decision.

This decision has profound organisational consequences. Infrastructure documentation no longer becomes outdated, because the system always matches what has been written. New environments — a second location, an additional test environment, disaster recovery — are not the product of days of configuration work, but of repeating existing code. Errors are caught before they occur, because changes are made explicit and can be reviewed.

IaC Overview

The decisions that come before the first line of code

Section titled “The decisions that come before the first line of code”

Before the first infrastructure description file is written, three organisational questions must be answered. They determine whether IaC becomes a stable operating model or a well-intentioned approach that isn’t followed in practice.

Who writes and maintains the infrastructure descriptions? The two extreme answers are both problematic. If only the Platform Team writes code, a bottleneck emerges: teams must wait for infrastructure instead of working independently. If every team acts fully autonomously without common standards, an incompatible tangle of approaches emerges. The methodically correct path lies in the middle: the Platform Team sets standards and provides reusable building blocks. Teams work autonomously within those boundaries.

How are changes to production environments approved? An infrastructure change in production made by one person, without review, without a record — that is the same risk as clicking manually in the console, only harder to undo. The sensible requirement is that production environment changes always run through an approved automated process, never directly by hand.

Where is the system state stored? IaC tools manage state: what is currently deployed? This state is critical to the functioning of the entire approach. It must be stored securely, centrally, and accessibly for the team — never on one person’s laptop, never in a publicly accessible system.

Infrastructure-as-Code changes the requirements for people who manage infrastructure. This is not primarily a technical challenge — it is a learning curve that requires time, practice, and a fault-tolerant learning environment.

What changes

Systems administrators who have managed infrastructure through configuration interfaces for years must develop a different mental model: instead of “I click Create,” they think “I describe the desired state.” This shift is learnable, but it takes time and must not be underestimated.

What stays

The infrastructure knowledge stays — and is the biggest advantage. Someone who understands how networks, security groups, and databases work learns the new description form quickly. Technical knowledge doesn’t need to be rebuilt, only the method of applying it.

What a good learning environment needs

Sandbox access for experimentation without consequences, pair work with more experienced colleagues, no punishment for mistakes in test environments. Someone who makes mistakes while learning, learns. Someone who makes mistakes and is criticised for them stops learning.

The Platform Team's role

The Platform Team is not a control body in this phase — it is an enabler. It provides learning materials, answers questions, reviews code alongside teams — not to evaluate, but to teach.

The transition to Infrastructure-as-Code is not a project with a start date and a sign-off. It is a maturity curve that unfolds over months and years. What matters is knowing the current position and taking the next sensible step — not trying to immediately reach the highest maturity level.

The first stage is the simplest: new resources are described as code, existing resources remain manual for now. This is incomplete, but a genuine start. Every new piece of infrastructure that emerges as code is a resource the team fully understands, has documented, and can restore.

The second stage brings all production resources under code control and introduces a central state store. Manual changes to production environments become the exception that must be documented.

The third stage connects IaC with automated quality and compliance checks. Changes are not only reviewed, but automatically evaluated against security and governance requirements before being executed.

Infrastructure-as-Code is not the goal — it is the means. The goal is a cloud environment that is traceable, secure, reproducible, and auditable. IaC is the most efficient path to get there.

Organisations that introduce IaC typically report three effects after six to twelve months: new environments emerge in hours, not days; infrastructure incidents are resolved faster because the path back to the baseline state is known; and audit requests are handled more efficiently because the documentation of state is created automatically.

These effects are more motivating than any abstract reference to best practices.

  1. Define the starting point: Which part of the infrastructure to begin with? New resources in a non-critical environment are the lowest-risk entry point. Existing complex production infrastructure is not a good starting point.

  2. Provide standards and building blocks: The Platform Team develops reusable descriptions for common infrastructure components. These building blocks are the lever that allows teams to start quickly without having to reinvent every step.

  3. Accompany the upskilling: Teams receive learning time, sandbox access, and support. This step is frequently underestimated — and is the most common reason why IaC introductions stall.

  4. Complete and evaluate the pilot: What worked? What was harder than expected? What needs to be adjusted? This evaluation is the foundation for the broader rollout.

  5. Expand incrementally: Based on pilot experience, bring in further teams until all production resources are under code control.

BASE

CI/CD Pipelines & Quality Gates

Build automated deployment pipelines with integrated, automatic quality and compliance checks before going live.

Adapting Operating ModelsCI/CD Pipelines In 1 trail

What CI/CD really means — beyond the tools

Section titled “What CI/CD really means — beyond the tools”

Continuous Integration and Continuous Deployment are not tools. They are organisational principles — a decision about how an organisation develops, validates, and operates software.

The core idea is simple: instead of deploying changes to production infrequently, manually, and riskily, they are delivered frequently, automatically, and in a controlled manner. Every change goes through the same quality checks. Every change is documented. Every change can be reversed.

CI/CD Overview

Organisations that take this path don’t deliver faster because they work less carefully. They deliver faster because they have automated the care.

Before a CI/CD pipeline is built, two cultural prerequisites must be in place. Without them, any technical implementation will fail against the organisation.

Tests are not optional additional work. A pipeline that checks automatically can only be as good as the tests it runs. If tests are seen as an annoying obligation to get through as quickly as possible, the pipeline becomes a formality — a green light that means nothing. Tests must be understood as an integral part of development work, not as overhead at the end.

Small, frequent changes rather than large, infrequent releases. The greatest risks in software changes come from the size of the change: the larger the package, the harder the debugging, the greater the impact when something goes wrong. CI/CD works best when teams learn to decompose changes into smaller, deliverable units. This is a way of working that must be learned — it is not natural for teams accustomed to monthly releases.

The decisions that come before implementation

Section titled “The decisions that come before implementation”

What blocks a change?

Quality gates in a pipeline are decisions, not technical defaults. What must pass before a change reaches production? Which tests are mandatory? Which security checks are non-negotiable? Setting these gates consciously — and documenting why they were set — is more important than the technical configuration.

Who is responsible for pipeline operations?

A pipeline is itself software — it must be maintained, updated, and analysed quickly when problems occur. If ownership is unclear, problems are ignored or bypassed. The Platform Team should own the pipeline infrastructure; individual teams own the tests and checks that their changes pass through.

How are failing pipelines handled?

How the organisation deals with red pipelines is a cultural question. Are failures fixed immediately? Or do they accumulate because “the test has been failing for weeks but it was never a real problem”? A failing pipeline that no longer alarms anyone has stopped providing safety.

What happens when something goes wrong?

A rollback process must be defined before it is needed — not in the moment when a critical problem has occurred. How long does a rollback take? Who can trigger one? Which teams must be informed? Answering these questions calmly in advance is much better than answering them under pressure.

The greatest danger at deployment is irreversibility: a change that causes a problem, and no fast path back. Deployment strategies address this through controlled introduction of changes.

The simplest strategy — rolling changes out progressively to a growing proportion of users — makes it possible to detect problems before all users are affected. If something goes wrong, five percent of users are affected, not a hundred. Rollback is a decision, not a catastrophe.

A more sophisticated strategy — running a new version in parallel with the old one until the new version has proven it works — eliminates the risk of outage entirely. If the new version shows problems, the switch is made back to the old one. No user experienced an outage.

Which strategy is right for which context depends on the criticality of the workload and the maturity of the team. A highly critical production environment needs different safety mechanisms than an internal tool.

Automated quality assurance as documentation

Section titled “Automated quality assurance as documentation”

An often-overlooked value of CI/CD is documentation. Every change that runs through a pipeline leaves evidence: what was changed, when, by whom, what checks were performed, was the result positive?

For organisations under regulatory requirements, this is not a side effect — it is a compliance asset. Instead of laboriously assembling audit requests from logs and memory, there is a complete, immutable trail for every production change. This substantially reduces the effort of audits.

Building delivery capability incrementally

Section titled “Building delivery capability incrementally”
  1. Choose a single pilot: Identify a workload where the team is motivated and willing to take risks. Not a core production service, but an important, non-critical one. Build the first pipeline together with the team — so the team understands and owns the pipeline, not just uses it.

  2. Define quality gates: What should this pipeline check? What does it block? These decisions are discussed with the team and documented. No automatic adoption of defaults — every gate has a reason.

  3. Evaluate the pilot phase: What worked well? What created friction? What should have been prevented? This review is the foundation for the broader rollout and should be conducted honestly.

  4. Rollout with support, not pressure: Bring in more teams — with support, not mandate. Teams that experience pipelines as a constraint build workarounds. Teams that experience pipelines as protection maintain them.

  5. Measure quality and learn: How long does a change take from development to production? How often does the pipeline fail and why? These metrics — not as a control mechanism, but as a learning source — show where investment should go.

The connection with Infrastructure-as-Code

Section titled “The connection with Infrastructure-as-Code”

CI/CD and Infrastructure-as-Code are not separate topics. In a mature environment, infrastructure changes are treated exactly like application changes: they run through a pipeline, are checked automatically, and changes to production environments happen exclusively through this approved process. This interplay — application code and infrastructure code in the same quality processes — is the hallmark of a mature cloud operating culture.

STEP

Landing Zone Governance

Provide standardized platform and application landing zones to inherit security requirements down to product teams.

Overview In 1 trail

A landing zone is the structured cloud foundation that defines how your organization operates on STACKIT from day 1. It combines governance, identity, security, network design, cost controls, and automation into one coherent baseline.

Without this foundation, migration waves typically stall due to missing approvals, inconsistent controls, and repeated platform decisions.

The following visual summarizes the core components that should be addressed for a reliable platform baseline. These six building blocks form a secure platform foundation:

  • Account Governance — project structure, hierarchy, and ownership model
  • Identity & Access Management — users, roles, and federated identity
  • Security & Compliance — guardrails, policies, and compliance controls
  • Network Architecture — segmentation, connectivity, and traffic control
  • Cost Management and Control — tagging, budgets, and spend visibility
  • Automation (IaC) — infrastructure as code, pipelines, and drift detection
Cloud Framework For more information: Landing Zones Overview Detailed guidance on platform and application landing zones, delivery model, and STACKIT acceleration assets. Open page
LIFT

Managed Service Provider (MSP) Governance

Enforce restrictive MSP policies, retain sovereignty over budget, IAM, and GDPR, and maintain internal runbooks to avoid partner lock-in.

Adapting Operating ModelsMSP Governance In 1 trail

Why MSP governance is often underestimated

Section titled “Why MSP governance is often underestimated”

Many organisations decide to outsource parts of their cloud operations to a managed service provider (MSP). The decision is often correct — operations become more reliable, and internal teams can focus on strategic topics.

What is frequently underestimated: governing the MSP relationship is itself a demanding task. Anyone who does not build a structured steering process gradually loses control of their own cloud environment — without noticing.

MSP Governance Overview

Operational tasks MSPs frequently assume:

  • Monitoring and alerting: 24/7 surveillance of infrastructure and applications
  • Incident response: first response and escalation for production incidents
  • Patch management: OS updates, security patches, Kubernetes upgrades
  • Backup management: backup verification, recovery tests, retention management
  • Compliance reporting: creation of audit evidence for regulatory reviews
  • Cost optimisation: monitoring for resource waste, rightsizing recommendations

What should remain internal:

  • Strategic cloud decisions (architecture, provider selection, investments)
  • Access management and IAM configuration (no root delegation to MSPs)
  • Data protection and GDPR responsibility (accountability cannot be delegated)
  • Budget responsibility and FinOps decisions
  • Crisis management and communication with regulators

An MSP contract must cover more than price and scope. Critical contract elements:

Service Level Agreement (SLA):

  • Availability SLA for the managed service (typical: 99.5% to 99.9%)
  • Response time by severity (P1: 15 minutes, P2: 1 hour, P3: 4 hours)
  • Resolution time targets with escalation path
  • SLA reporting frequency and format

Penalties and credits: Without economic consequences for SLA violations, SLAs have little steering effect. Typical: credits of 10–30% of the monthly service fee for demonstrated SLA shortfall.

Subcontractor provisions: Which tasks may the MSP further delegate? Every sub-delegation must be transparent and must meet the same data protection and security standards — relevant for GDPR Art. 28 para. 4.

Exit provisions: How is the handover to another provider or back to the internal team governed? Timelines, knowledge transfer obligations, documentation handover, access return.

The most common governance error: the MSP has too many rights, for too long.

Best practice for MSP access:

  • No permanent admin rights — instead just-in-time access for maintenance windows
  • Dedicated MSP accounts (no use of employee accounts)
  • All MSP activities visible in the audit log
  • Monthly review of active MSP permissions
  • Immediate deactivation at contract end

The internal IAM team (or CCoE) retains owner rights on all projects. The MSP operates with editor rights in defined scopes — never with owner rights.

MSP governance requires structure. Without regular review, the relationship drifts in a direction that the client does not notice — until there is a problem.

MSP lock-in is often not contractual but knowledge-based: the internal team no longer knows how its own infrastructure functions.

Documentation requirements:

  • All architecture decisions documented in writing, maintained in the internal wiki
  • Runbooks for all operational processes — ownership lies internally, not with the MSP
  • Change log for all configuration changes, maintained by MSP and accessible internally
  • Quarterly knowledge transfer sessions: MSP explains to the internal team what has changed

STACKIT experience

Does the MSP have demonstrable experience with STACKIT? Are there reference customers from similar industries? STACKIT-certified partners provide a structured entry point.

Compliance expertise

Does the MSP understand the regulatory requirements of your sector? GDPR, TISAX, BAIT, KRITIS — an MSP without compliance expertise is unsuitable for regulated environments.

Transparent processes

Can the MSP present their incident response process, change management procedures and security concept? Providers who evade these questions are a risk.

Exit readiness

A reputable MSP actively shapes the exit clause — because it knows that a good relationship is the best client retention. Anyone who blocks exit provisions is creating dependency.

  1. Define scope — What is outsourced, what remains internal? In writing, as the basis for the tender.

  2. MSP tender or selection — At least three offers, a structured evaluation matrix, reference conversations with existing clients.

  3. Contract negotiation — SLAs, penalties, subcontractor provisions, exit clause, GDPR data processing agreement. Legal counsel specialised in cloud contracts is recommended.

  4. Implement access concept — Set up MSP accounts, define permission scope, activate audit logging, establish review process.

  5. Establish review rhythm — Weekly, monthly and quarterly reviews anchored in the calendar. Create agenda templates.

  6. Agree documentation requirements in writing — What does the MSP deliver, in what format, at what interval? As a contractual component, not a verbal assurance.

The SLA requirements for MSPs are derived from the DR tier classifications of the affected workloads. The internal Service Catalogue defines which services IT delivers itself and which the MSP assumes.

GOAL

Disaster Recovery and Business Continuity

Proactively define, document, and regularly simulate RTO and RPO targets for business-critical systems on STACKIT.

Adapting Operating ModelsDisaster Recovery In 1 trail

Disaster recovery is the discipline nobody practises — until they have to. Then every decision that could have been made earlier costs money and trust.

For most organisations, cloud migration is the moment when DR requirements are defined in writing for the first time. This is the opportunity: clearly, proportionately, and tested.

RTO and RPO — the two fundamental questions

Section titled “RTO and RPO — the two fundamental questions”

Every system has an implicit recovery time and an acceptable data loss threshold. Most organisations do not know these figures — and pay the price when an incident occurs.

Recovery Time Objective (RTO): How long may a system be unavailable after an outage? An RTO of 4 hours means: after at most 4 hours, the system must be functioning again.

Recovery Point Objective (RPO): How much data loss is acceptable? An RPO of 1 hour means: at most the last 60 minutes of data may be lost.

The important rule: The more aggressive the RTO and RPO, the higher the ongoing costs. An RTO of 15 minutes with an RPO of 5 minutes requires hot-standby environments. An RTO of 24 hours with an RPO of 12 hours can be met with daily backups and manual processes.

The right answer is: proportionate, not maximalist.

Dr Overview

Tier 1 — Mission Critical (RTO < 1h, RPO < 15min)

Section titled “Tier 1 — Mission Critical (RTO < 1h, RPO < 15min)”

Systems whose failure immediately leads to revenue loss, regulatory consequences, or safety risks. Examples: payment systems, core banking systems, industrial production control.

Tier 2 — Business Critical (RTO < 4h, RPO < 1h)

Section titled “Tier 2 — Business Critical (RTO < 4h, RPO < 1h)”

Systems that must be restored within a few hours, but whose outage is tolerable in the short term. Examples: ERP systems, CRM, internal portals.

Tier 3 — Business Important (RTO < 24h, RPO < 4h)

Section titled “Tier 3 — Business Important (RTO < 24h, RPO < 4h)”

Systems that can be restored within one working day. Examples: reporting systems, secondary databases, near-test systems.

Tier 4 — Standard (RTO < 72h, RPO < 24h)

Section titled “Tier 4 — Standard (RTO < 72h, RPO < 24h)”

Systems without immediate business criticality. Examples: development environments, archive systems, secondary analytics tools.

DR and compliance — what regulators expect

Section titled “DR and compliance — what regulators expect”

For regulated industries, DR is not a voluntary best practice but a regulatory requirement.

BAIT (banking): Requires documented contingency plans, regular tests, and failure scenarios for all material IT systems. Recovery targets must be approved by the management board.

KRITIS (critical infrastructure): Obliges KRITIS operators to business continuity management in accordance with BSI IT-Grundschutz or ISO 22301. DR tests are mandatory; results must be documented.

GDPR: Art. 32 requires the ability to restore the availability of personal data after an incident. An RPO of more than 24 hours for GDPR-relevant systems is difficult to justify.

DORA (Digital Operational Resilience Act, from January 2025): Obliges financial institutions to comprehensive ICT risk management with explicit BCM requirements and mandatory resilience tests.

STACKIT as a partner: STACKIT operates data centres in Germany with ISO 27001 and SOC 2 certification. Data residency in de-01/de-02 ensures that backup data does not leave Germany — a decisive advantage for regulated industries.

DR tests — why they rarely happen and how to change that

Section titled “DR tests — why they rarely happen and how to change that”

The most common DR mistake is never testing the plan. DR tests are regularly postponed because the production environment is considered too critical to risk a simulated failover, no clear ownership for DR tests exists, and tests are seen as effort without direct benefit.

The solution: Anchor DR tests as a regular engineering ritual, not as an exceptional event.

Four test formats in ascending intensity:

1. Tabletop exercise (annually, low effort): The team works through a failure scenario together — on the whiteboard. No systems touched. Identifies gaps in processes and accountabilities.

2. Component test (quarterly): A specific backup or failover mechanism is tested in the staging environment. Duration: a few hours.

3. Partial failover test (semi-annually): A non-production-critical system is actually switched to the standby site. Measures actual RTO. Duration: half a day.

4. Full DR exercise (annually): Complete simulation of a data centre outage for all Tier 1 and Tier 2 systems. Requires planning and alignment with stakeholders. Delivers the real evidence for regulators.

CIO / IT Lead

Approves RTO/RPO targets, acts as escalation authority in an emergency, signs off DR tests and reports results to the board and regulators.

Platform Team

Implements and maintains DR infrastructure, conducts component tests, documents recovery procedures.

Application owners

Classify their systems into tier classes, define application-specific recovery steps, participate in tabletop exercises.

STACKIT

Guarantees availability of cloud infrastructure in accordance with SLA, provides zone-to-zone replication, delivers audit logs for regulatory evidence.

  1. Inventory and classify all systems — Tier assignment for each workload, documented and confirmed by the business unit.

  2. Have RTO/RPO targets approved by the board or CIO — Not as an IT decision, but as a business decision with cost implications.

  3. Implement DR architecture per tier — From simple backups in object storage to hot standby in the second STACKIT zone.

  4. Document recovery procedures — As a runbook, not a concept. Concrete steps that can be executed in an emergency without further questions.

  5. Conduct the first tabletop exercise — Within 30 days of go-live. Output: a list of open items with owner and deadline.

  6. Establish and maintain a DR test calendar — Quarterly component tests, semi-annual partial failover, annual full exercise.

Trail historyAdded Sep 10, 2026?UpdatedNo updates · 1 bar = 1 week i
Maintainers
  • ?Name not public?Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site.
??Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site.Contributed in STACKIT
  • Name not public?Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site. · Sep 10, 2026