Restructure teams
The model of stream-aligned teams with genuine operational responsibility — and how a platform engineering team enables all others without becoming a bottleneck.
Last updated on
A migration only unfolds its full impact once operating and process structures modernize: from ITIL to You-Build-It-You-Run-It structures to DORA metrics.
Introduce the cloud-native operating model and analyze the financial risk of failing to transform.
Cloud technology does not automatically change how your organisation develops and operates software. Many organisations have the same long deployment cycles after migration, the same silos between development and operations, the same manual processes — just now in the cloud.
This chapter addresses the organisational changes that make cloud effective: how teams are structured, how responsibility is distributed, how change management and incident response are adapted to cloud speed — and how to measure whether the transformation has actually taken place.
Restructure teams
The model of stream-aligned teams with genuine operational responsibility — and how a platform engineering team enables all others without becoming a bottleneck.
Infrastructure as Code
Why clicking in the console must no longer be the standard — and how Infrastructure as Code delivers reproducibility, security and speed simultaneously.
Adapt ITIL to cloud speed
Which change management processes must work differently in the cloud — and how to define Standard Changes so that compliance and agility are not opposites.
Measure success
The four DORA metrics as an objective measure of delivery performance — with benchmarks and a realistic improvement roadmap.
A financial services provider migrated 40 workloads to STACKIT in eight months. Technically a success. Organisationally a sobering experience.
Deployment cycles: still six weeks. Not because the technology was too slow — but because the change management process continued to route every patch through a three-person CAB committee.
Manual configuration: still via the console. Not because no IaC tool was available — but because nobody had learned to use it, and nobody was given the time.
Costs: 60 % higher than budgeted. Not because STACKIT is expensive — but because no team took ownership of its cloud cost share.
EUR 420,000 in annual additional costs. The migration had succeeded. The transformation had failed.
Cloud transformation creates new roles and changes existing ones. System administrators become platform engineers. Infrastructure teams become DevOps teams. Some roles that exist today will no longer be needed in three years — at least not in their current form.
Communicating this reality openly and actively shaping it — with clear qualification paths, fair transition processes, and the works council as a partner — is not only morally required. It is strategically necessary. Anyone who does not have these conversations will lose exactly the people who are most urgently needed for the transformation.
The DevOps & YBIYRI chapter addresses this with concrete team topologies and role transitions.
Adapting operating models presupposes Cloud Empowerment — teams need the capability to practise new ways of working.
Roll out the You Build It, You Run It principle in stages, pilot team, expansion, full adoption, to avoid operational bottlenecks.
In the classic IT model there is a structural handover: developers build an application and pass it to operations. Operations deploys, monitors and repairs.
This handover systematically creates three problems:
Friction and waiting times: Every change passes through a handover process. What is technically finished waits for the capacity of the operations team.
Unclear accountability: When an application fails in production, the first question is: is it the code or the infrastructure? That question costs time — which is especially valuable during an incident.
Different priorities: Development teams want to deliver new features. Operations teams want stability. These interests are structurally in conflict as long as they sit in different teams.
YBIYRI (You Build It You Run It) resolves all three problems through a simple shift: the team that builds an application also operates it in production. Full responsibility, full competence.
Stream-aligned teams are fully responsible for a service or product — from development through to production operations:
This full ownership is the core of YBIYRI. It changes the mindset: a person who knows they will be woken at night if the application fails builds it differently.
The platform engineering team (often evolved from the CCoE) is not an ops team in the classical sense. It builds and operates the internal platform that makes life easier for all stream-aligned teams:
The platform team is not an approval body — it is an enabler. Stream-aligned teams should be able to get their work done without raising a ticket for every infrastructure topic.
YBIYRI is attractive as a concept — but it fails when the organisational prerequisites are absent.
Good observability
No team can operate a service it cannot see. Before YBIYRI is introduced, the observability platform must be in place: dashboards, alerts, runbooks.
Automated deployments
On-call teams cannot deploy manually at 2am. Automated, reproducible deployments via CI/CD are a prerequisite — not a nice-to-have.
Clear Service Level Objectives
SLOs define when an incident is genuinely critical. Without this clarity, every minor anomaly wakes someone up. With SLOs, escalation only happens when a target is at risk.
Psychological safety
Teams do not take on genuine responsibility when mistakes are punished. A culture in which incidents are treated as learning opportunities (blameless post-mortem) is a fundamental prerequisite.
Fair on-call compensation
YBIYRI without fair compensation for on-call duty leads to burnout and the loss of the best staff. On-call must be clearly regulated and appropriately compensated — before introduction, not after.
Sufficient team size
A team of three people cannot sustain a healthy on-call rotation. YBIYRI requires that teams are large enough (typically 5–8 people) to distribute on-call duty.
SLOs (Service Level Objectives) are a team’s written commitment to its service. They answer: what is “good enough” for us — and from when do we escalate?
A complete SLO defines at least three dimensions:
Availability: What proportion of the time must the service be available? An SLO of 99.9% means: a maximum of 8.7 hours of downtime per year is accepted.
Latency: How quickly must the service respond? “95% of all requests under 200 milliseconds” is a concrete, measurable target.
Error rate: What proportion of requests may be answered with an error? “Less than 0.1% 5xx responses” protects users from systemic problems.
These three numbers together define the team’s quality commitment. They are the basis for on-call decisions: escalation happens when an SLO is at risk — not at every anomaly.
Pilot with one team (months 1–3): A volunteer team adopts YBIYRI for a non-critical service. Set up on-call rotation, define SLOs, gather first experiences. Document lessons learned.
Expansion (months 3–6): Three to five further teams adopt YBIYRI. The platform team delivers self-service tooling that reduces cognitive load. Shared runbooks and playbooks are created.
Full adoption (months 6–12): All new services are built under the YBIYRI model. The classic ops team gradually transforms into platform engineering. Legacy systems remain in the classic model during the transition.
This question occupies leadership more than any technical question. The honest answer: the classic ops team does not disappear. It transforms.
Staff with strong infrastructure knowledge are highly valuable in the platform engineering team: they know operational problems, they understand what can go wrong in production, they have the experience that developers often lack.
The qualification measures for this transition are described in the Workforce Transition chapter. What must be communicated early: nobody loses their job through YBIYRI — but the tasks change.
Coordinate through decentralized dependency boards and establish an inner source principle for sharing Terraform and Helm modules.
Technically, a cloud migration is plannable. Organisationally, it rarely is. As soon as multiple teams work on the same platform in parallel, dependencies emerge that nobody fully sees at the start.
A common pattern: the application team waits for the platform team’s landing zone. The platform team waits for the security team’s IAM decision. The security team waits for the regulatory sign-off from compliance. Everyone is waiting — and the go-live date approaches.
This chapter describes how dependencies become visible early and how teams can work in parallel despite dependencies.
In classical IT projects, there is a project manager who knows and steers all dependencies. In cloud transformations, work is distributed across teams with different cadences, different priorities, and different incentive structures.
A platform team works in two-week sprints. A compliance team works in quarterly cycles. An application team works according to a product roadmap. These three rhythms systematically produce misalignment.
The solution is not stronger central control — it is better visibility and clearer interfaces.
Before teams work together, it should be clear: what is the relationship between them?
Platform Team → application teams: The Platform Team provides services (Kubernetes clusters, networking, IAM templates). Application teams consume these services. Interaction should run as much as possible through self-service interfaces (service catalogue, Terraform modules, documentation) — not through tickets.
CCoE → all teams: The CCoE sets standards, reviews architecture decisions, and supports on escalations. It is not an approval committee — it is an enabler with veto rights on critical security and compliance questions.
Application teams with each other: Shared-service dependencies (e.g. a central database used by multiple teams) are the most common coordination bottleneck. These should be explicitly treated as a “platform service” and operated by the responsible team as an internal product with an SLA.
A simple but effective instrument: a board (physical or digital) that visualises all cross-team dependencies.
| Dependency | Delivering team | Receiving team | Planned by | Status |
|---|---|---|---|---|
| Landing Zone Prod | Platform | App Team Billing | Week 14 | in progress |
| IAM role concept | Security | Platform | Week 12 | blocked |
| Data protection approval | Compliance | App Team CRM | Week 16 | pending |
| CI/CD Pipeline Template | Platform | App Team Logistics | Week 13 | ready |
The dependency board is updated weekly. Blockages become visible immediately — before they become a bottleneck.
Daily standup (team-internal)
15 minutes daily. What was done yesterday? What is planned today? What is blocking? Blockages with cross-team dependencies are escalated immediately — not as a ticket, but as direct contact.
Platform sync (weekly)
45 minutes. Platform Team presents: what is newly available? What is coming in the next two weeks? Which changes have breaking-change potential? All application teams are represented.
Dependency review (fortnightly)
30 minutes. Review of the dependency board. Which dependencies are open, which are escalating? Decisions on shifts are made here, not via email.
Steering committee (monthly)
Management, CIO, tech leads. Status of the overall migration, strategic course adjustments, resource decisions. No detailed discussions — only decisions.
When all teams use the same STACKIT infrastructure, a shared code-base effect emerges: Terraform modules, Helm charts, and CI/CD templates are developed twice.
Inner Source means: platform components are shared internally like open-source projects. Every team can contribute. The Platform Team maintains and reviews. Result: no duplicates, faster iteration, shared quality awareness.
In practice: an internal git repository with Terraform modules for STACKIT resources, versioned and documented. Application teams use, improve, and share back.
Sometimes processes are not enough. Then clear escalation paths are needed.
Escalation trigger: A cross-team dependency has been blocked for two weeks and has put a go-live date at risk.
Escalation path:
What should never happen: Resolving blockages through email ping-pong without a defined escalation point. Time pressure does not solve structural dependency problems.
Document team topology — Who works with whom? In what mode (collaboration, X-as-a-Service, facilitating)? In writing, as the basis for all other coordination formats.
Set up the dependency board — Start simply: a table in Confluence or Jira is sufficient. Important: a weekly review date in the calendar.
Establish sync formats — Create platform sync and dependency review as recurring appointments. Create agenda templates.
Communicate the escalation path — All teams know who to escalate to and when. Not as a threat, but as a safety net.
Create an Inner Source repository — Start with the Terraform module that most teams need. Document how contributions are made.
Adapt classic ITIL processes, for example reducing the CMDB to strategic information and establishing IaC as the source of configuration truth.
ITIL (IT Infrastructure Library) was the response to the chaos of early IT organisations: standardised processes for change management, incident management, problem management and service design. For most organisations, ITIL represented a necessary step towards maturity.
In cloud environments, ITIL comes under pressure. A Change Advisory Board that meets once a week is not compatible with an organisation aiming for hundreds of deployments per day. A 14-day change request process blocks the agile iteration that the cloud is designed to enable.
The answer is not: abolish ITIL. The answer is: evolve ITIL.
What remains from ITIL:
What changes fundamentally:
Classic change management was designed for a world in which every change to production systems was manual, risky, and difficult to reverse. In that world, a formal approval process made sense.
In cloud environments with Infrastructure-as-Code and automated tests, the risk assessment changes fundamentally. A Terraform change that has been automatically checked against policies before deployment, tested in staging, and reviewed by two people via pull request has a different risk profile than a manual configuration change at night.
The modernised change model:
Standard Changes (frequent, well-documented, low risk) are fully automated. No CAB, no change ticket — the automation is the approval process. Examples: container image updates, configuration changes for known parameters, resource scaling.
Normal Changes (changes to critical system components) go through a pull request process with two reviewers, an automated test suite, and a deployment window. The “CAB” is the asynchronous review by experienced colleagues — faster, more efficient, with the same level of assurance.
Emergency Changes (critical production fixes) have an accelerated process with documentation completed afterwards. They are discussed in the post-incident review: was the emergency procedure correctly applied? What prevents the same issue from being treated as an emergency again?
The core model of incident management — detection, triage, escalation, resolution, post-mortem — is just as valid in cloud environments as in classical IT environments.
What changes is the speed and the expectations.
Detection: Classically through user reports or manual monitoring. Today through automated observability systems that detect anomalies before users are affected.
Triage: Classically through manual diagnosis via SSH and log files. Today through central dashboards, distributed tracing and structured log search.
On-call structure: Cloud environments require a 24/7 on-call rotation. This is culturally the biggest shift for many IT organisations: who is on call, how is it compensated, how is burnout prevented?
Post-mortem culture: Blameless post-mortems are standard in DevOps cultures — no blame assignment, but systemic learning. For ITIL-shaped organisations, this is often a cultural shift that requires time and leadership commitment.
The Configuration Management Database (CMDB) was ITIL’s answer to the question: what is running where, in what configuration, with what dependencies on what else? In classical environments it was valuable — and notoriously difficult to keep current.
In cloud environments with Infrastructure-as-Code, the CMDB is no longer a primary system. Infrastructure-as-Code (Terraform, Ansible) is the new CMDB — with the decisive advantage that it is automatically correct: what is in Terraform is (after the last apply) reality.
What this means for the organisation:
SLAs with business units remain important — but their design changes. Classical SLAs spoke about availability in percentages on a monthly basis. Cloud SLAs can be more granular.
Availability SLA
Monthly availability of the service. Based on the STACKIT platform SLA, reduced by an internal operational buffer. Transparently communicated and documented in the service catalogue.
Response Time SLA
Response times by severity. P1 within 15 minutes, P2 within 1 hour — not as a promise, but as a measurable commitment with monthly reporting.
Recovery Time SLA
RTO and RPO per tier class (Tier 1 to 4). Not as a theoretical figure, but as a regularly tested and demonstrated capability.
Change Lead Time
How long does it take to get an approved change into production? For standard changes: hours. For normal changes: days. Not weeks.
Abolishing ITIL processes overnight is just as wrong as transferring them unchanged into the cloud. The right approach is gradual.
Inventory: Which ITIL processes exist? Which of them make sense in the cloud, and which create drag?
Identify quick wins: Automating standard changes is typically fast to implement and immediately noticeable.
Revise the change process: Increase CAB frequency or replace with asynchronous review. Introduce an automation gateway as an alternative to manual approvals.
Introduce blameless post-mortems: Culturally demanding, but decisive for continuous improvement. Leadership must model that mistakes are learning opportunities.
Adapt CMDB strategy: Establish IaC as the primary configuration source; reduce the CMDB to strategic information.
Introduce the four DORA metrics to objectively and regularly measure, and improve, software delivery performance.
DORA (DevOps Research and Assessment) found in a seven-year study of more than 32,000 organisations: four metrics correlate strongly with business outcomes including market share, profitability and employee satisfaction.
DORA metrics measure outcomes, not effort. An organisation that improves these four metrics objectively improves its ability to deliver software reliably and quickly — regardless of which methods or tools it uses.
This metric measures delivery capability. Teams with high deployment frequency deliver small, targeted changes — which reduces risk and shortens feedback cycles.
| Performance level | Deployment frequency | Meaning |
|---|---|---|
| Elite | Multiple times daily | Continuous deployment, smallest possible batches |
| High | Once daily to once weekly | Regular, structured releases |
| Medium | Once weekly to once monthly | Still significant manual effort |
| Low | Less than monthly | High release overhead, concentrated risk |
How it is measured: Number of deployments to production per time unit — from the CI/CD system or deployment tracking.
This metric measures the efficiency of the delivery process. A long lead time means many manual steps, waiting times, approvals or heavyweight tests stand between a code change and the production environment.
| Performance level | Lead time | Meaning |
|---|---|---|
| Elite | Less than 1 hour | Fully automated pipeline |
| High | 1 day to 1 week | Good degree of automation |
| Medium | 1 week to 1 month | Many manual processes |
| Low | More than 1 month | Heavyweight release processes |
This metric measures quality. A high change failure rate means production deployments regularly cause problems — despite tests and reviews.
| Performance level | Change failure rate |
|---|---|
| Elite | 0–15% |
| High | 16–30% |
| Medium/Low | 31–60% |
A high change failure rate is often the driver of low deployment frequency: when every second deployment breaks something, deployments become rarer and riskier.
This metric measures resilience. Incidents will always occur — the ability to respond and recover quickly distinguishes elite organisations from others.
| Performance level | MTTR |
|---|---|
| Elite | Less than 1 hour |
| High | Less than 1 day |
| Medium | 1 day to 1 week |
| Low | More than 1 week |
Most organisations beginning their first cloud migration start at Low to Medium level. This is not a criticism — it is the natural starting point of an organisation coming from classic IT operating models.
| Metric | Typical starting point | DORA level | Year 1 target | Year 2 target |
|---|---|---|---|---|
| Deployment Frequency | 1× every 2–4 weeks | Low | 1× / week (Medium) | 1× / day (High) |
| Lead Time | 3–6 weeks | Low | 1–5 days | under 1 day |
| Change Failure Rate | 30–50% | Low/Medium | under 20% | under 10% |
| MTTR | 2–5 days | Low | under 4 hours | under 1 hour |
DORA metrics are not abstract concepts — they need concrete data sources:
Deployment Frequency
Data source: CI/CD system (GitHub Actions, GitLab CI, Jenkins). Every successful production deployment is counted. Straightforward to automate.
Lead Time
Data source: Git timestamp of the commit combined with deployment timestamp from CI/CD. The difference is the lead time. DORA tools such as Sleuth or LinearB automate this calculation.
Change Failure Rate
Data source: Incidents from incident tracking (PagerDuty, OpsGenie, JIRA) combined with deployments. Proportion of deployments that led to an incident within 24 hours.
MTTR
Data source: Incident tracking system. Time between incident opening and incident closure for production incidents.
DORA metrics in management reporting: These four numbers belong in the monthly IT report to leadership — as an objective measure of delivery capability, not as technical metrics for developers. They are as relevant to the business as NPS or conversion rate.
Quarter 1 — build the foundation: Two to three pilot services receive a complete CI/CD pipeline. Automated testing is introduced. The first DORA dashboard is built. Deployment Frequency: Low → Medium.
Quarter 2 — expand automation: Standard changes are fully automated. The observability stack is complete. Lead time reduces through fewer manual steps. Lead Time: Low → High.
Quarter 3 — improve quality: Test coverage is expanded. Security scanning is embedded in CI/CD. Incident response process is optimised. Change Failure Rate and MTTR improve.
Quarter 4 — scale: DORA practices are rolled out to all teams. DORA metrics are part of the monthly management report. All four metrics in the High range.
DORA metrics — in particular MTTR and Change Failure Rate — cannot be measured reliably without a good Observability setup. An organisation that does not detect incidents within minutes cannot reduce MTTR to under one hour. Building observability is therefore not a technical side matter — it is the prerequisite for measurable DORA improvement.