OPS 1. How do you share operational responsibility between those who build and those who run?
Last updated on
This question is about people, which is why it is easy to skip and expensive to get wrong. Every other question in this pillar describes a practice. This one describes the conditions under which those practices are adopted at all.
The observable symptom of getting it wrong is a system that is pleasant to develop and miserable to operate: error messages that identify nothing, configuration that requires tribal knowledge, failure modes that produce a page and no diagnostic path. Nobody designs that deliberately. It is what happens when the feedback loop between building and running is missing.
Best practices
Section titled “Best practices”OPS 1.1Give the team that builds a system a genuine stake in running itOPS 1.2Make incident review blameless in practice, not only in policyOPS 1.3Define ownership so that every component has a name against itOPS 1.4Fund operational work explicitly rather than expecting it to fit in the gaps
OPS 1.1 Give the team that builds a system a genuine stake in running it
Section titled “OPS 1.1 Give the team that builds a system a genuine stake in running it”Risk if not established: Medium
The mechanism is a feedback loop, not a staffing model. Engineers who experience the consequences of their design choices make different choices, and no amount of guidance substitutes for that.
There are several ways to close it, and carrying a pager is only the most obvious. Rotating developers through operational duty, having the building team own the service level objectives, or simply requiring that they attend their own incident reviews all create the same feedback with different intensity. Choose the one your organization can actually sustain, because a rotation that burns people out closes the loop once and then breaks it permanently.
Watch for the failure mode where responsibility moves without authority. A team that is woken by incidents but cannot change the architecture, the dependencies or the priorities has been given the cost of ownership without the means to reduce it. That produces resentment rather than better systems.
Where a separate operations function exists, the loop can still be closed through shared objectives, joint incident review, and a hard rule that the operating team can refuse to accept a system that is not operable. The rule matters more than the structure.
On STACKIT. No platform feature applies. This is organizational design.
The one platform-adjacent decision is the resource hierarchy, because it determines whether
ownership can be expressed at all. See OPS 1.3.
Tradeoffs. Cost Optimization. Operational duty consumes engineering capacity that would otherwise ship features, and the return arrives as incidents that did not happen, which is invisible in every report. This is the compounding argument from the Operational Excellence tradeoffs, and it is the hardest one to win with a spreadsheet.
Verify. Who is called when your most critical flow breaks at three in the morning, and did that person write any of it? If not, what feedback do the people who wrote it receive?
OPS 1.2 Make incident review blameless in practice, not only in policy
Section titled “OPS 1.2 Make incident review blameless in practice, not only in policy”Risk if not established: Medium
The useful information about an incident is held by the person closest to it, and it is only available if giving it is safe. A team that fears the consequences of an incident will optimize for not being implicated in one, which means fewer deployments, less experimentation, and incident reports that omit the interesting part.
Blamelessness is a property of behaviour rather than of a stated policy, and it is tested at the first incident with an obvious human cause. What happens then is what everyone will remember.
Two concrete tests. Does a review ever conclude with a person’s name as the cause? That is the cheapest available answer and it stops short of the useful one, which is why the system allowed a normal human error to have that consequence. And are near misses reported? People only report the incidents that nobody would have noticed when doing so is safe, and near misses are where the cheapest learning is.
This does not mean the absence of accountability. Teams are accountable for the systems they run and for acting on what reviews find. What is removed is individual blame for the honest mistakes that any competent person would eventually make.
On STACKIT. No platform feature applies.
One adjacent point worth stating: audit logs record who performed an action, and their legitimate
purposes are investigation and compliance rather than attribution of blame. How that record is
used after an incident is a cultural decision, and using it to find someone to blame will end the
reporting of near misses immediately. See OPS 9.3.
Tradeoffs. None material. The resistance is cultural rather than economic, and it usually comes from a belief that consequences drive care. They drive concealment.
Verify. Read your last three incident reviews. How many identify a person as the cause, and how many identify why the system permitted the error to have that effect?
OPS 1.3 Define ownership so that every component has a name against it
Section titled “OPS 1.3 Define ownership so that every component has a name against it”Risk if not established: Medium
Unowned components are where reliability goes to decay. Nobody patches them, nobody notices when their monitoring breaks, and during an incident the first fifteen minutes are spent finding out whose they are.
Ownership needs to be at team level rather than individual level, so that it survives people changing roles. It needs to cover everything, including the shared infrastructure, the internal tooling and the pipeline. And it needs to be discoverable without asking anyone, which means it lives somewhere queryable rather than in institutional memory.
The uncomfortable part of the exercise is the components nobody claims. Those are usually shared services that grew organically, and they are disproportionately represented in incidents. Assigning them is unpopular and it is the point.
On STACKIT. The Resource Manager hierarchy of organization, folders and projects is where ownership becomes structural rather than documented. When projects follow ownership, then access, billing and the resource inventory all align with the team responsible, and the ownership question answers itself.
That alignment is worth deciding early because retrofitting it means moving resources between
projects. It is the same structure SEC 2 asks for on blast-radius grounds and COST 2 asks for
on attribution grounds, which means one decision serves three pillars.
IAM roles and memberships assigned per project make the owning team explicit in the access model rather than only in a wiki.
Tradeoffs. Cost Optimization. A hierarchy that mirrors ownership can prevent resource sharing that would have been cheaper. That is usually the right trade, and it is a trade.
Verify. Pick three resources at random from your environment. How long does it take to establish which team owns each, and where did you find the answer?
OPS 1.4 Fund operational work explicitly rather than expecting it to fit in the gaps
Section titled “OPS 1.4 Fund operational work explicitly rather than expecting it to fit in the gaps”Risk if not established: Medium
Automation, tooling, runbooks, rehearsals and incident actions all cost engineering time and produce nothing a customer sees. When they are expected to happen in whatever time is left over, they do not happen, because there is never time left over.
The consequence compounds in a specific way: the team with no time to automate has less time next
quarter, because it is doing manually what the automation would have done. That is the toil spiral
described in OPS 10.
Make it a stated allocation rather than an aspiration. A proportion of capacity reserved for operational work, protected in the same way feature commitments are protected, is the only version that survives a delivery deadline.
Incident actions deserve their own treatment, because they arrive unplanned and compete with planned work. An action from an incident review that has no owner and no date will not be done, and the review that produced it was theatre.
On STACKIT. No platform feature applies.
What the platform does affect is how much operational work exists. Choosing a managed service over
a self-operated equivalent moves work to the provider, which is the honest way to reduce the
allocation rather than pretending it is smaller. That comparison belongs in the cost model under
COST 1, where operational labour is routinely omitted.
Tradeoffs. Cost Optimization. Reserved capacity for operational work is capacity not spent on features, and the benefit is deferred and diffuse. The argument is about the slope of the next two years rather than this quarter.
Verify. What proportion of your team’s capacity went to operational work last quarter, and was that a decision or a residual? How many actions from incident reviews are still open?
Related
Section titled “Related”OPS 9Incident management, which depends on the culture this question establishesOPS 10Toil elimination, which is what happens when operational work is fundedSEC 2Segmentation andCOST 2Cost attribution, which need the same resource hierarchyREL 1Reliability targets, which need an owner to survive cost pressure- Operational Excellence tradeoffs