Skip to content
Beta

Target Operating Model & Process Agility

Last updated on

Stackit LogoStackit Logo
STACKIT

Target Operating Model & Process Agility

A migration only unfolds its full impact once operating and process structures modernize: from ITIL to You-Build-It-You-Run-It structures to DORA metrics.

PLAN

Target Operating Model Shift

Introduce the cloud-native operating model and analyze the financial risk of failing to transform.

Overview In 1 trail

Cloud technology does not automatically change how your organisation develops and operates software. Many organisations have the same long deployment cycles after migration, the same silos between development and operations, the same manual processes — just now in the cloud.

This chapter addresses the organisational changes that make cloud effective: how teams are structured, how responsibility is distributed, how change management and incident response are adapted to cloud speed — and how to measure whether the transformation has actually taken place.

Restructure teams

The model of stream-aligned teams with genuine operational responsibility — and how a platform engineering team enables all others without becoming a bottleneck.

Infrastructure as Code

Why clicking in the console must no longer be the standard — and how Infrastructure as Code delivers reproducibility, security and speed simultaneously.

Adapt ITIL to cloud speed

Which change management processes must work differently in the cloud — and how to define Standard Changes so that compliance and agility are not opposites.

Measure success

The four DORA metrics as an objective measure of delivery performance — with benchmarks and a realistic improvement roadmap.

Operating Model Overview

When the operating model is not transformed

Section titled “When the operating model is not transformed”

A financial services provider migrated 40 workloads to STACKIT in eight months. Technically a success. Organisationally a sobering experience.

Deployment cycles: still six weeks. Not because the technology was too slow — but because the change management process continued to route every patch through a three-person CAB committee.

Manual configuration: still via the console. Not because no IaC tool was available — but because nobody had learned to use it, and nobody was given the time.

Costs: 60 % higher than budgeted. Not because STACKIT is expensive — but because no team took ownership of its cloud cost share.

EUR 420,000 in annual additional costs. The migration had succeeded. The transformation had failed.

Organisational workforce transition: the forgotten task

Section titled “Organisational workforce transition: the forgotten task”

Cloud transformation creates new roles and changes existing ones. System administrators become platform engineers. Infrastructure teams become DevOps teams. Some roles that exist today will no longer be needed in three years — at least not in their current form.

Communicating this reality openly and actively shaping it — with clear qualification paths, fair transition processes, and the works council as a partner — is not only morally required. It is strategically necessary. Anyone who does not have these conversations will lose exactly the people who are most urgently needed for the transformation.

The DevOps & YBIYRI chapter addresses this with concrete team topologies and role transitions.

Adapting operating models presupposes Cloud Empowerment — teams need the capability to practise new ways of working.

  1. Clarify team structure with DevOps & YBIYRI — before technical details are decided.
  2. Introduce Infrastructure as Code as the standard — not as an option.
  3. Automate delivery with CI/CD pipelines — bound to compliance gates.
  4. Adapt ITIL to cloud reality — without abandoning compliance.
  5. Measure progress with DORA Metrics — objectively and regularly.
BASE

DevOps Topologies & YBIYRI Alignment

Roll out the You Build It, You Run It principle in stages, pilot team, expansion, full adoption, to avoid operational bottlenecks.

Adapting Operating ModelsDevOps & YBIYRI In 1 trail

In the classic IT model there is a structural handover: developers build an application and pass it to operations. Operations deploys, monitors and repairs.

This handover systematically creates three problems:

Friction and waiting times: Every change passes through a handover process. What is technically finished waits for the capacity of the operations team.

Unclear accountability: When an application fails in production, the first question is: is it the code or the infrastructure? That question costs time — which is especially valuable during an incident.

Different priorities: Development teams want to deliver new features. Operations teams want stability. These interests are structurally in conflict as long as they sit in different teams.

YBIYRI (You Build It You Run It) resolves all three problems through a simple shift: the team that builds an application also operates it in production. Full responsibility, full competence.

YBIYRI Overview

Stream-aligned teams are fully responsible for a service or product — from development through to production operations:

  • They develop, deploy and operate their application
  • They have their own budget and a dedicated STACKIT project
  • They carry on-call duty for their service
  • They deploy themselves — without a central ops team as gatekeeper
  • They define their own SLOs and measure themselves against them

This full ownership is the core of YBIYRI. It changes the mindset: a person who knows they will be woken at night if the application fails builds it differently.

The platform engineering team (often evolved from the CCoE) is not an ops team in the classical sense. It builds and operates the internal platform that makes life easier for all stream-aligned teams:

  • Landing zone, network infrastructure, Kubernetes clusters
  • CI/CD toolchain and deployment pipeline templates
  • Shared Terraform modules, validated and maintained
  • Observability platform (monitoring, logging, alerting)
  • Internal documentation and architecture reviews

The platform team is not an approval body — it is an enabler. Stream-aligned teams should be able to get their work done without raising a ticket for every infrastructure topic.

YBIYRI is attractive as a concept — but it fails when the organisational prerequisites are absent.

Good observability

No team can operate a service it cannot see. Before YBIYRI is introduced, the observability platform must be in place: dashboards, alerts, runbooks.

Automated deployments

On-call teams cannot deploy manually at 2am. Automated, reproducible deployments via CI/CD are a prerequisite — not a nice-to-have.

Clear Service Level Objectives

SLOs define when an incident is genuinely critical. Without this clarity, every minor anomaly wakes someone up. With SLOs, escalation only happens when a target is at risk.

Psychological safety

Teams do not take on genuine responsibility when mistakes are punished. A culture in which incidents are treated as learning opportunities (blameless post-mortem) is a fundamental prerequisite.

Fair on-call compensation

YBIYRI without fair compensation for on-call duty leads to burnout and the loss of the best staff. On-call must be clearly regulated and appropriately compensated — before introduction, not after.

Sufficient team size

A team of three people cannot sustain a healthy on-call rotation. YBIYRI requires that teams are large enough (typically 5–8 people) to distribute on-call duty.

Service Level Objectives — what teams own

Section titled “Service Level Objectives — what teams own”

SLOs (Service Level Objectives) are a team’s written commitment to its service. They answer: what is “good enough” for us — and from when do we escalate?

A complete SLO defines at least three dimensions:

Availability: What proportion of the time must the service be available? An SLO of 99.9% means: a maximum of 8.7 hours of downtime per year is accepted.

Latency: How quickly must the service respond? “95% of all requests under 200 milliseconds” is a concrete, measurable target.

Error rate: What proportion of requests may be answered with an error? “Less than 0.1% 5xx responses” protects users from systemic problems.

These three numbers together define the team’s quality commitment. They are the basis for on-call decisions: escalation happens when an SLO is at risk — not at every anomaly.

  1. Pilot with one team (months 1–3): A volunteer team adopts YBIYRI for a non-critical service. Set up on-call rotation, define SLOs, gather first experiences. Document lessons learned.

  2. Expansion (months 3–6): Three to five further teams adopt YBIYRI. The platform team delivers self-service tooling that reduces cognitive load. Shared runbooks and playbooks are created.

  3. Full adoption (months 6–12): All new services are built under the YBIYRI model. The classic ops team gradually transforms into platform engineering. Legacy systems remain in the classic model during the transition.

This question occupies leadership more than any technical question. The honest answer: the classic ops team does not disappear. It transforms.

Staff with strong infrastructure knowledge are highly valuable in the platform engineering team: they know operational problems, they understand what can go wrong in production, they have the experience that developers often lack.

The qualification measures for this transition are described in the Workforce Transition chapter. What must be communicated early: nobody loses their job through YBIYRI — but the tasks change.

STEP

Multi-Team Coordination & Inner Source

Coordinate through decentralized dependency boards and establish an inner source principle for sharing Terraform and Helm modules.

Adapting Operating ModelsMulti-Team Coordination In 1 trail

The coordination problem in cloud projects

Section titled “The coordination problem in cloud projects”

Technically, a cloud migration is plannable. Organisationally, it rarely is. As soon as multiple teams work on the same platform in parallel, dependencies emerge that nobody fully sees at the start.

A common pattern: the application team waits for the platform team’s landing zone. The platform team waits for the security team’s IAM decision. The security team waits for the regulatory sign-off from compliance. Everyone is waiting — and the go-live date approaches.

This chapter describes how dependencies become visible early and how teams can work in parallel despite dependencies.

The topology problem: why classical project structures fail

Section titled “The topology problem: why classical project structures fail”

In classical IT projects, there is a project manager who knows and steers all dependencies. In cloud transformations, work is distributed across teams with different cadences, different priorities, and different incentive structures.

A platform team works in two-week sprints. A compliance team works in quarterly cycles. An application team works according to a product roadmap. These three rhythms systematically produce misalignment.

The solution is not stronger central control — it is better visibility and clearer interfaces.

1. Define team topology and interaction modes

Section titled “1. Define team topology and interaction modes”

Before teams work together, it should be clear: what is the relationship between them?

Platform Team → application teams: The Platform Team provides services (Kubernetes clusters, networking, IAM templates). Application teams consume these services. Interaction should run as much as possible through self-service interfaces (service catalogue, Terraform modules, documentation) — not through tickets.

CCoE → all teams: The CCoE sets standards, reviews architecture decisions, and supports on escalations. It is not an approval committee — it is an enabler with veto rights on critical security and compliance questions.

Application teams with each other: Shared-service dependencies (e.g. a central database used by multiple teams) are the most common coordination bottleneck. These should be explicitly treated as a “platform service” and operated by the responsible team as an internal product with an SLA.

A simple but effective instrument: a board (physical or digital) that visualises all cross-team dependencies.

The dependency board is updated weekly. Blockages become visible immediately — before they become a bottleneck.

Daily standup (team-internal)

15 minutes daily. What was done yesterday? What is planned today? What is blocking? Blockages with cross-team dependencies are escalated immediately — not as a ticket, but as direct contact.

Platform sync (weekly)

45 minutes. Platform Team presents: what is newly available? What is coming in the next two weeks? Which changes have breaking-change potential? All application teams are represented.

Dependency review (fortnightly)

30 minutes. Review of the dependency board. Which dependencies are open, which are escalating? Decisions on shifts are made here, not via email.

Steering committee (monthly)

Management, CIO, tech leads. Status of the overall migration, strategic course adjustments, resource decisions. No detailed discussions — only decisions.

4. Inner Source principle for platform components

Section titled “4. Inner Source principle for platform components”

When all teams use the same STACKIT infrastructure, a shared code-base effect emerges: Terraform modules, Helm charts, and CI/CD templates are developed twice.

Inner Source means: platform components are shared internally like open-source projects. Every team can contribute. The Platform Team maintains and reviews. Result: no duplicates, faster iteration, shared quality awareness.

In practice: an internal git repository with Terraform modules for STACKIT resources, versioned and documented. Application teams use, improve, and share back.

Sometimes processes are not enough. Then clear escalation paths are needed.

Escalation trigger: A cross-team dependency has been blocked for two weeks and has put a go-live date at risk.

Escalation path:

  1. Direct conversation between the affected team leads: 24-hour deadline for resolution
  2. If no result: involvement of the CCoE as neutral mediator
  3. If no result: steering committee makes the decision — with all consequences

What should never happen: Resolving blockages through email ping-pong without a defined escalation point. Time pressure does not solve structural dependency problems.

  1. Document team topology — Who works with whom? In what mode (collaboration, X-as-a-Service, facilitating)? In writing, as the basis for all other coordination formats.

  2. Set up the dependency board — Start simply: a table in Confluence or Jira is sufficient. Important: a weekly review date in the calendar.

  3. Establish sync formats — Create platform sync and dependency review as recurring appointments. Create agenda templates.

  4. Communicate the escalation path — All teams know who to escalate to and when. Not as a threat, but as a safety net.

  5. Create an Inner Source repository — Start with the Terraform module that most teams need. Document how contributions are made.

LIFT

ITIL Modernisation for Cloud Speed

Adapt classic ITIL processes, for example reducing the CMDB to strategic information and establishing IaC as the source of configuration truth.

Adapting Operating ModelsITIL Modernisation In 1 trail

ITIL (IT Infrastructure Library) was the response to the chaos of early IT organisations: standardised processes for change management, incident management, problem management and service design. For most organisations, ITIL represented a necessary step towards maturity.

In cloud environments, ITIL comes under pressure. A Change Advisory Board that meets once a week is not compatible with an organisation aiming for hundreds of deployments per day. A 14-day change request process blocks the agile iteration that the cloud is designed to enable.

The answer is not: abolish ITIL. The answer is: evolve ITIL.

ITIL Overview

What remains from ITIL:

  • The core principles: structured processes, clear accountabilities, documentation, continuous improvement
  • Incident Management: structured response to outages, escalation paths, post-incident reviews
  • Problem Management: root cause analysis, prevention of recurrence
  • Service Level Management: agreements on availability and quality with business units

What changes fundamentally:

  • Change Management: from manual CAB to automated pipeline gate
  • Release Management: from quarterly releases to continuous deployment
  • Configuration Management: from CMDB as a manual database to Infrastructure-as-Code as the single source of truth

Change management — the biggest transformation

Section titled “Change management — the biggest transformation”

Classic change management was designed for a world in which every change to production systems was manual, risky, and difficult to reverse. In that world, a formal approval process made sense.

In cloud environments with Infrastructure-as-Code and automated tests, the risk assessment changes fundamentally. A Terraform change that has been automatically checked against policies before deployment, tested in staging, and reviewed by two people via pull request has a different risk profile than a manual configuration change at night.

The modernised change model:

Standard Changes (frequent, well-documented, low risk) are fully automated. No CAB, no change ticket — the automation is the approval process. Examples: container image updates, configuration changes for known parameters, resource scaling.

Normal Changes (changes to critical system components) go through a pull request process with two reviewers, an automated test suite, and a deployment window. The “CAB” is the asynchronous review by experienced colleagues — faster, more efficient, with the same level of assurance.

Emergency Changes (critical production fixes) have an accelerated process with documentation completed afterwards. They are discussed in the post-incident review: was the emergency procedure correctly applied? What prevents the same issue from being treated as an emergency again?

Incident management — what stays and what modernises

Section titled “Incident management — what stays and what modernises”

The core model of incident management — detection, triage, escalation, resolution, post-mortem — is just as valid in cloud environments as in classical IT environments.

What changes is the speed and the expectations.

Detection: Classically through user reports or manual monitoring. Today through automated observability systems that detect anomalies before users are affected.

Triage: Classically through manual diagnosis via SSH and log files. Today through central dashboards, distributed tracing and structured log search.

On-call structure: Cloud environments require a 24/7 on-call rotation. This is culturally the biggest shift for many IT organisations: who is on call, how is it compensated, how is burnout prevented?

Post-mortem culture: Blameless post-mortems are standard in DevOps cultures — no blame assignment, but systemic learning. For ITIL-shaped organisations, this is often a cultural shift that requires time and leadership commitment.

The Configuration Management Database (CMDB) was ITIL’s answer to the question: what is running where, in what configuration, with what dependencies on what else? In classical environments it was valuable — and notoriously difficult to keep current.

In cloud environments with Infrastructure-as-Code, the CMDB is no longer a primary system. Infrastructure-as-Code (Terraform, Ansible) is the new CMDB — with the decisive advantage that it is automatically correct: what is in Terraform is (after the last apply) reality.

What this means for the organisation:

  • CMDB updates are no longer a manual process — they arise automatically from the IaC workflow
  • The CMDB can focus on higher-value information: business context, licence management, lifecycle management
  • Teams must understand that IaC is their new “source of truth” — and treat it accordingly

Service level management in cloud environments

Section titled “Service level management in cloud environments”

SLAs with business units remain important — but their design changes. Classical SLAs spoke about availability in percentages on a monthly basis. Cloud SLAs can be more granular.

Availability SLA

Monthly availability of the service. Based on the STACKIT platform SLA, reduced by an internal operational buffer. Transparently communicated and documented in the service catalogue.

Response Time SLA

Response times by severity. P1 within 15 minutes, P2 within 1 hour — not as a promise, but as a measurable commitment with monthly reporting.

Recovery Time SLA

RTO and RPO per tier class (Tier 1 to 4). Not as a theoretical figure, but as a regularly tested and demonstrated capability.

Change Lead Time

How long does it take to get an approved change into production? For standard changes: hours. For normal changes: days. Not weeks.

Practical recommendation: stepwise migration

Section titled “Practical recommendation: stepwise migration”

Abolishing ITIL processes overnight is just as wrong as transferring them unchanged into the cloud. The right approach is gradual.

  1. Inventory: Which ITIL processes exist? Which of them make sense in the cloud, and which create drag?

  2. Identify quick wins: Automating standard changes is typically fast to implement and immediately noticeable.

  3. Revise the change process: Increase CAB frequency or replace with asynchronous review. Introduce an automation gateway as an alternative to manual approvals.

  4. Introduce blameless post-mortems: Culturally demanding, but decisive for continuous improvement. Leadership must model that mistakes are learning opportunities.

  5. Adapt CMDB strategy: Establish IaC as the primary configuration source; reduce the CMDB to strategic information.

GOAL

DORA Metrics & Performance Measurement

Introduce the four DORA metrics to objectively and regularly measure, and improve, software delivery performance.

Adapting Operating ModelsDORA Metrics In 1 trail

DORA (DevOps Research and Assessment) found in a seven-year study of more than 32,000 organisations: four metrics correlate strongly with business outcomes including market share, profitability and employee satisfaction.

DORA metrics measure outcomes, not effort. An organisation that improves these four metrics objectively improves its ability to deliver software reliably and quickly — regardless of which methods or tools it uses.

Dora Overview

1. Deployment Frequency — how often does the team deploy to production?

Section titled “1. Deployment Frequency — how often does the team deploy to production?”

This metric measures delivery capability. Teams with high deployment frequency deliver small, targeted changes — which reduces risk and shortens feedback cycles.

How it is measured: Number of deployments to production per time unit — from the CI/CD system or deployment tracking.

2. Lead Time for Changes — time from commit to production deployment

Section titled “2. Lead Time for Changes — time from commit to production deployment”

This metric measures the efficiency of the delivery process. A long lead time means many manual steps, waiting times, approvals or heavyweight tests stand between a code change and the production environment.

3. Change Failure Rate — proportion of deployments that cause an incident

Section titled “3. Change Failure Rate — proportion of deployments that cause an incident”

This metric measures quality. A high change failure rate means production deployments regularly cause problems — despite tests and reviews.

A high change failure rate is often the driver of low deployment frequency: when every second deployment breaks something, deployments become rarer and riskier.

4. Mean Time to Restore (MTTR) — time to resolve a production incident

Section titled “4. Mean Time to Restore (MTTR) — time to resolve a production incident”

This metric measures resilience. Incidents will always occur — the ability to respond and recover quickly distinguishes elite organisations from others.

Most organisations beginning their first cloud migration start at Low to Medium level. This is not a criticism — it is the natural starting point of an organisation coming from classic IT operating models.

DORA metrics are not abstract concepts — they need concrete data sources:

Deployment Frequency

Data source: CI/CD system (GitHub Actions, GitLab CI, Jenkins). Every successful production deployment is counted. Straightforward to automate.

Lead Time

Data source: Git timestamp of the commit combined with deployment timestamp from CI/CD. The difference is the lead time. DORA tools such as Sleuth or LinearB automate this calculation.

Change Failure Rate

Data source: Incidents from incident tracking (PagerDuty, OpsGenie, JIRA) combined with deployments. Proportion of deployments that led to an incident within 24 hours.

MTTR

Data source: Incident tracking system. Time between incident opening and incident closure for production incidents.

DORA metrics in management reporting: These four numbers belong in the monthly IT report to leadership — as an objective measure of delivery capability, not as technical metrics for developers. They are as relevant to the business as NPS or conversion rate.

Quarter 1 — build the foundation: Two to three pilot services receive a complete CI/CD pipeline. Automated testing is introduced. The first DORA dashboard is built. Deployment Frequency: Low → Medium.

Quarter 2 — expand automation: Standard changes are fully automated. The observability stack is complete. Lead time reduces through fewer manual steps. Lead Time: Low → High.

Quarter 3 — improve quality: Test coverage is expanded. Security scanning is embedded in CI/CD. Incident response process is optimised. Change Failure Rate and MTTR improve.

Quarter 4 — scale: DORA practices are rolled out to all teams. DORA metrics are part of the monthly management report. All four metrics in the High range.

DORA metrics — in particular MTTR and Change Failure Rate — cannot be measured reliably without a good Observability setup. An organisation that does not detect incidents within minutes cannot reduce MTTR to under one hour. Building observability is therefore not a technical side matter — it is the prerequisite for measurable DORA improvement.

Trail historyAdded Sep 10, 2026?UpdatedNo updates · 1 bar = 1 week i
Maintainers
  • ?Name not public?Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site.
??Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site.Contributed in STACKIT
  • Name not public?Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site. · Sep 10, 2026