---
id: OPS01
pillar: operational-excellence
title: OPS 1. How do you share operational responsibility between those who build and those who run?
description: When builders are insulated from operating what they built, the system becomes convenient to write and unpleasant to run. How to close that feedback loop.
status: draft
services: []
sidebar:
  order: 10
  label: Shared ownership
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/ops-01-shared-ownership/"
source_file: "docs/architecture/pillars/operational-excellence/ops-01-shared-ownership.mdx"
---

This question is about people, which is why it is easy to skip and expensive to get wrong. Every
other question in this pillar describes a practice. This one describes the conditions under which
those practices are adopted at all.

The observable symptom of getting it wrong is a system that is pleasant to develop and miserable
to operate: error messages that identify nothing, configuration that requires tribal knowledge,
failure modes that produce a page and no diagnostic path. Nobody designs that deliberately. It is
what happens when the feedback loop between building and running is missing.

## Best practices

- [`OPS 1.1`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-11-give-the-team-that-builds-a-system-a-genuine-stake-in-running-it) Give the team that builds a system a genuine stake in running it
- [`OPS 1.2`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-12-make-incident-review-blameless-in-practice-not-only-in-policy) Make incident review blameless in practice, not only in policy
- [`OPS 1.3`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-13-define-ownership-so-that-every-component-has-a-name-against-it) Define ownership so that every component has a name against it
- [`OPS 1.4`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-14-fund-operational-work-explicitly-rather-than-expecting-it-to-fit-in-the-gaps) Fund operational work explicitly rather than expecting it to fit in the gaps

---

## OPS 1.1 Give the team that builds a system a genuine stake in running it

**Risk if not established:** Medium

The mechanism is a feedback loop, not a staffing model. Engineers who experience the consequences
of their design choices make different choices, and no amount of guidance substitutes for that.

There are several ways to close it, and carrying a pager is only the most obvious. Rotating
developers through operational duty, having the building team own the service level objectives, or
simply requiring that they attend their own incident reviews all create the same feedback with
different intensity. Choose the one your organization can actually sustain, because a rotation
that burns people out closes the loop once and then breaks it permanently.

Watch for the failure mode where responsibility moves without authority. A team that is woken by
incidents but cannot change the architecture, the dependencies or the priorities has been given
the cost of ownership without the means to reduce it. That produces resentment rather than better
systems.

Where a separate operations function exists, the loop can still be closed through shared
objectives, joint incident review, and a hard rule that the operating team can refuse to accept a
system that is not operable. The rule matters more than the structure.

**On STACKIT.** No platform feature applies. This is organizational design.

The one platform-adjacent decision is the resource hierarchy, because it determines whether
ownership can be expressed at all. See [`OPS 1.3`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-13-define-ownership-so-that-every-component-has-a-name-against-it).

**Tradeoffs.** **Cost Optimization.** Operational duty consumes engineering capacity that would
otherwise ship features, and the return arrives as incidents that did not happen, which is
invisible in every report. This is the compounding argument from the
[Operational Excellence tradeoffs](/architecture/pillars/operational-excellence/tradeoffs/), and it is the hardest one to win with a
spreadsheet.

**Verify.** Who is called when your most critical flow breaks at three in the morning, and did
that person write any of it? If not, what feedback do the people who wrote it receive?

---

## OPS 1.2 Make incident review blameless in practice, not only in policy

**Risk if not established:** Medium

The useful information about an incident is held by the person closest to it, and it is only
available if giving it is safe. A team that fears the consequences of an incident will optimize
for not being implicated in one, which means fewer deployments, less experimentation, and incident
reports that omit the interesting part.

Blamelessness is a property of behaviour rather than of a stated policy, and it is tested at the
first incident with an obvious human cause. What happens then is what everyone will remember.

Two concrete tests. Does a review ever conclude with a person's name as the cause? That is the
cheapest available answer and it stops short of the useful one, which is why the system allowed a
normal human error to have that consequence. And are near misses reported? People only report the
incidents that nobody would have noticed when doing so is safe, and near misses are where the
cheapest learning is.

This does not mean the absence of accountability. Teams are accountable for the systems they run
and for acting on what reviews find. What is removed is individual blame for the honest mistakes
that any competent person would eventually make.

**On STACKIT.** No platform feature applies.

One adjacent point worth stating: audit logs record who performed an action, and their legitimate
purposes are investigation and compliance rather than attribution of blame. How that record is
used after an incident is a cultural decision, and using it to find someone to blame will end the
reporting of near misses immediately. See [`OPS 9.3`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-93-review-for-structural-causes-rather-than-proximate-ones).

**Tradeoffs.** None material. The resistance is cultural rather than economic, and it usually
comes from a belief that consequences drive care. They drive concealment.

**Verify.** Read your last three incident reviews. How many identify a person as the cause, and
how many identify why the system permitted the error to have that effect?

---

## OPS 1.3 Define ownership so that every component has a name against it

**Risk if not established:** Medium

Unowned components are where reliability goes to decay. Nobody patches them, nobody notices when
their monitoring breaks, and during an incident the first fifteen minutes are spent finding out
whose they are.

Ownership needs to be at team level rather than individual level, so that it survives people
changing roles. It needs to cover everything, including the shared infrastructure, the internal
tooling and the pipeline. And it needs to be discoverable without asking anyone, which means it
lives somewhere queryable rather than in institutional memory.

The uncomfortable part of the exercise is the components nobody claims. Those are usually shared
services that grew organically, and they are disproportionately represented in incidents.
Assigning them is unpopular and it is the point.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/resource-manager/">Resource Manager</LinkChip>
hierarchy of
organization, folders and projects is where ownership becomes structural rather than documented.
When projects follow ownership, then access, billing and the resource inventory all align with the
team responsible, and the ownership question answers itself.

That alignment is worth deciding early because retrofitting it means moving resources between
projects. It is the same structure [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) asks for on blast-radius grounds and [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) asks for
on attribution grounds, which means one decision serves three pillars.

<LinkChip href="https://docs.stackit.cloud/platform/access-and-identity/roles-permissions/roles-permissions/">IAM roles and memberships</LinkChip>
assigned per project make the
owning team explicit in the access model rather than only in a wiki.

**Tradeoffs.** **Cost Optimization.** A hierarchy that mirrors ownership can prevent resource
sharing that would have been cheaper. That is usually the right trade, and it is a trade.

**Verify.** Pick three resources at random from your environment. How long does it take to
establish which team owns each, and where did you find the answer?

---

## OPS 1.4 Fund operational work explicitly rather than expecting it to fit in the gaps

**Risk if not established:** Medium

Automation, tooling, runbooks, rehearsals and incident actions all cost engineering time and
produce nothing a customer sees. When they are expected to happen in whatever time is left over,
they do not happen, because there is never time left over.

The consequence compounds in a specific way: the team with no time to automate has less time next
quarter, because it is doing manually what the automation would have done. That is the toil spiral
described in [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/).

Make it a stated allocation rather than an aspiration. A proportion of capacity reserved for
operational work, protected in the same way feature commitments are protected, is the only version
that survives a delivery deadline.

Incident actions deserve their own treatment, because they arrive unplanned and compete with
planned work. An action from an incident review that has no owner and no date will not be done,
and the review that produced it was theatre.

**On STACKIT.** No platform feature applies.

What the platform does affect is how much operational work exists. Choosing a managed service over
a self-operated equivalent moves work to the provider, which is the honest way to reduce the
allocation rather than pretending it is smaller. That comparison belongs in the cost model under
[`COST 1`](/architecture/pillars/cost-optimization/cost-01-cost-model/), where operational labour is routinely omitted.

**Tradeoffs.** **Cost Optimization.** Reserved capacity for operational work is capacity not
spent on features, and the benefit is deferred and diffuse. The argument is about the slope of the
next two years rather than this quarter.

**Verify.** What proportion of your team's capacity went to operational work last quarter, and was
that a decision or a residual? How many actions from incident reviews are still open?

---

## Related

- [`OPS 9`](/architecture/pillars/operational-excellence/ops-09-incident-management/) Incident management, which depends on the culture this question establishes
- [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/) Toil elimination, which is what happens when operational work is funded
- [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) Segmentation and [`COST 2`](/architecture/pillars/cost-optimization/cost-02-cost-attribution/) Cost attribution, which need the same resource hierarchy
- [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) Reliability targets, which need an owner to survive cost pressure
- <LinkChip href="/architecture/pillars/operational-excellence/tradeoffs/">Operational Excellence tradeoffs</LinkChip>
