Good observability
No team can operate a service it cannot see. Before YBIYRI is introduced, the observability platform must be in place: dashboards, alerts, runbooks.
In the classic IT model there is a structural handover: developers build an application and pass it to operations. Operations deploys, monitors and repairs.
This handover systematically creates three problems:
Friction and waiting times: Every change passes through a handover process. What is technically finished waits for the capacity of the operations team.
Unclear accountability: When an application fails in production, the first question is: is it the code or the infrastructure? That question costs time — which is especially valuable during an incident.
Different priorities: Development teams want to deliver new features. Operations teams want stability. These interests are structurally in conflict as long as they sit in different teams.
YBIYRI (You Build It You Run It) resolves all three problems through a simple shift: the team that builds an application also operates it in production. Full responsibility, full competence.
Stream-aligned teams are fully responsible for a service or product — from development through to production operations:
This full ownership is the core of YBIYRI. It changes the mindset: a person who knows they will be woken at night if the application fails builds it differently.
The platform engineering team (often evolved from the CCoE) is not an ops team in the classical sense. It builds and operates the internal platform that makes life easier for all stream-aligned teams:
The platform team is not an approval body — it is an enabler. Stream-aligned teams should be able to get their work done without raising a ticket for every infrastructure topic.
YBIYRI is attractive as a concept — but it fails when the organisational prerequisites are absent.
Good observability
No team can operate a service it cannot see. Before YBIYRI is introduced, the observability platform must be in place: dashboards, alerts, runbooks.
Automated deployments
On-call teams cannot deploy manually at 2am. Automated, reproducible deployments via CI/CD are a prerequisite — not a nice-to-have.
Clear Service Level Objectives
SLOs define when an incident is genuinely critical. Without this clarity, every minor anomaly wakes someone up. With SLOs, escalation only happens when a target is at risk.
Psychological safety
Teams do not take on genuine responsibility when mistakes are punished. A culture in which incidents are treated as learning opportunities (blameless post-mortem) is a fundamental prerequisite.
Fair on-call compensation
YBIYRI without fair compensation for on-call duty leads to burnout and the loss of the best staff. On-call must be clearly regulated and appropriately compensated — before introduction, not after.
Sufficient team size
A team of three people cannot sustain a healthy on-call rotation. YBIYRI requires that teams are large enough (typically 5–8 people) to distribute on-call duty.
SLOs (Service Level Objectives) are a team’s written commitment to its service. They answer: what is “good enough” for us — and from when do we escalate?
A complete SLO defines at least three dimensions:
Availability: What proportion of the time must the service be available? An SLO of 99.9% means: a maximum of 8.7 hours of downtime per year is accepted.
Latency: How quickly must the service respond? “95% of all requests under 200 milliseconds” is a concrete, measurable target.
Error rate: What proportion of requests may be answered with an error? “Less than 0.1% 5xx responses” protects users from systemic problems.
These three numbers together define the team’s quality commitment. They are the basis for on-call decisions: escalation happens when an SLO is at risk — not at every anomaly.
Pilot with one team (months 1–3): A volunteer team adopts YBIYRI for a non-critical service. Set up on-call rotation, define SLOs, gather first experiences. Document lessons learned.
Expansion (months 3–6): Three to five further teams adopt YBIYRI. The platform team delivers self-service tooling that reduces cognitive load. Shared runbooks and playbooks are created.
Full adoption (months 6–12): All new services are built under the YBIYRI model. The classic ops team gradually transforms into platform engineering. Legacy systems remain in the classic model during the transition.
This question occupies leadership more than any technical question. The honest answer: the classic ops team does not disappear. It transforms.
Staff with strong infrastructure knowledge are highly valuable in the platform engineering team: they know operational problems, they understand what can go wrong in production, they have the experience that developers often lack.
The qualification measures for this transition are described in the Workforce Transition chapter. What must be communicated early: nobody loses their job through YBIYRI — but the tasks change.