OPS 10. How do you find and eliminate toil?
Last updated on
Toil is manual work that recurs, produces no lasting value, and scales with the size of the system. Applying a patch by hand, clearing a stuck queue every Tuesday, provisioning an account on request, copying figures into a spreadsheet before a meeting.
It is dangerous precisely because it looks like productivity. It is visible, people thank you for it, and it feels like the job. Meanwhile it consumes the capacity that would have removed it, which is why teams with the most toil have the least time to address it.
Best practices
Section titled “Best practices”OPS 10.1Measure recurring manual work rather than absorbing itOPS 10.2Automate what recurs, starting with the riskiest rather than the most frequentOPS 10.3Remove the need rather than automating the workaroundOPS 10.4Budget for elimination explicitly
OPS 10.1 Measure recurring manual work rather than absorbing it
Section titled “OPS 10.1 Measure recurring manual work rather than absorbing it”Risk if not established: Medium
Toil is invisible in every reporting system because nobody records it. It happens between the tickets, and the team absorbs it until the absorption fails.
Measure it for a bounded period rather than forever. Two weeks of everyone noting recurring manual operations and roughly how long each took produces a list that is uncomfortable and actionable. Precision is not the point; the ranking is.
Capture three things per item: how often it recurs, how long it takes, and what happens if it is done wrong. The third is what separates the annoying from the dangerous, and it is usually the one nobody has considered.
Watch for the work that is invisible even to the people doing it: the mental overhead of remembering that something must be checked, the interruptions, the context switching. Those cost more than the minutes suggest and they are the reason a small recurring task can dominate a week.
On STACKIT. No platform feature measures your manual work.
Two sources give partial visibility. The
audit log records actions taken through the
Portal, the CLI and the API, so a repeated manual action against platform resources is visible
there, within the retention window OPS 7.4 covers. And a high ratio of console actions to pipeline actions is itself
a signal, both of toil and of the drift OPS 3.3 looks for.
Tradeoffs. Operational Excellence. Two weeks of light record-keeping. The main resistance is that the exercise makes visible how much time goes to work nobody planned, which is the point and is uncomfortable.
Verify. List the recurring manual operations your team performed last month, with frequency and duration. What proportion of the team’s time do they represent?
OPS 10.2 Automate what recurs, starting with the riskiest rather than the most frequent
Section titled “OPS 10.2 Automate what recurs, starting with the riskiest rather than the most frequent”Risk if not established: Medium
The instinct is to automate the most frequent task, because the time saved is easiest to calculate. The better first target is usually the one where a mistake is most expensive.
A weekly task that saves twenty minutes is worth automating. A quarterly task that touches production data, has eleven steps, and has gone wrong twice is worth automating first, because the value is in eliminating the error rather than the minutes.
Automation should fail safely and be observable. A script that stops and alerts is better than one
that continues past an unexpected state, and one that reports what it did is what makes it
trustworthy enough to leave alone. Automation with a bad condition applies its mistake everywhere
at machine speed, which is the point OPS 5.3 also makes.
Not everything should be automated. Work that is genuinely different each time, needs judgement,
or happens twice a year and takes ten minutes is better documented than automated, since the
automation would itself become a thing to maintain. OPS 8 is the alternative, and choosing
between the two deliberately is part of this.
On STACKIT. Automation Service is the managed platform for scheduled and triggered automation across STACKIT services, with templates for recurring patterns.
For virtual machines, Run Command executes scripts across servers with command templates , which covers a large share of the repetitive operational work on Compute Engine.
Patching is one category worth naming, because it is high-frequency and high-consequence. Server
Update Management
provides scheduled operating system updates for Linux and Windows, which removes a recurring
manual operation and satisfies SEC 8 at the same time.
Beyond those, the CLI ,
API and
SDKs mean anything
reachable through the platform can be scripted, which is the same reach OPS 3 depends on.
Tradeoffs. Operational Excellence, against itself: automation is code that must be maintained, tested and understood. Automating something that changes frequently produces a maintenance burden larger than the toil it replaced.
Verify. Of the recurring operations you measured, which have been automated in the last six months? Which was chosen first, and was the reason frequency or consequence?
OPS 10.3 Remove the need rather than automating the workaround
Section titled “OPS 10.3 Remove the need rather than automating the workaround”Risk if not established: Medium
Automating a recurring task makes it cheaper. Removing the reason it recurs makes it free, and the second is available more often than teams assume.
A queue that has to be cleared weekly is telling you about a design problem. A service that must be restarted every few days has a leak. A certificate that needs manual renewal has an automation gap upstream. Automating any of those hides the signal and makes the underlying defect permanent, because nobody feels it any more.
Ask why before asking how. If a task recurs because of a design decision, changing the design removes it entirely. If it recurs because of a missing capability, building that capability removes a class of tasks rather than one.
The pattern to look for is toil that keeps reappearing in the same component. OPS 3.4 makes the
same observation about emergency changes: a component that repeatedly needs manual intervention is
describing itself.
Sometimes the answer is to delete the thing. A report nobody reads, an environment nobody opens, a
job whose output goes nowhere. That is SUS 6 and it is the cheapest fix available.
On STACKIT. No platform feature does this. It is a design decision each time.
The platform-shaped version of the question is worth asking though: some toil exists because
something is self-operated that could be consumed as a managed service. Moving a self-run database
to PostgreSQL Flex or a self-run
CI system to STACKIT
Pipelines
removes the operational work rather than automating it. The comparison belongs in COST 1, which
routinely omits the labour on the self-operated side.
Tradeoffs. Cost Optimization. Removing a need is usually a larger change than automating a task, and it competes with feature work on a longer timescale. It is also the only option that scales.
Verify. For your three largest sources of toil, why does each recur? For how many is the answer a design decision that could be changed rather than a task that could be scripted?
OPS 10.4 Budget for elimination explicitly
Section titled “OPS 10.4 Budget for elimination explicitly”Risk if not established: Medium
Toil elimination has the worst possible shape for getting funded: the cost is immediate and visible, the benefit is deferred and diffuse, and the work produces nothing a customer sees.
Left to compete with features on merit, it loses every time, and the loss compounds. Each quarter
of unaddressed toil consumes more of the following quarter, which is the spiral described in
OPS 1.4.
Reserve capacity rather than intending to find it. A stated proportion, protected in the same way feature commitments are protected, is the only version that survives a deadline. The number matters less than the fact that it is a decision.
Set a ceiling as a trigger. Where toil exceeds an agreed share of capacity, elimination takes priority over new work until it is back under. That converts a judgement call into a rule, which is what makes it survive the quarter where everything is urgent.
Report the result. Toil eliminated is capacity returned, and it is the only form in which this work shows up as a number anyone outside the team recognizes.
On STACKIT. No platform feature applies.
The one platform decision that changes the arithmetic is managed versus self-operated, since it
moves operational work off your team permanently. That comparison is only honest if the labour is
counted, which is COST 1 and the reason a managed database that looks expensive next to a
virtual machine is frequently cheaper.
Tradeoffs. Cost Optimization. Reserved capacity is capacity not spent on features. The
argument is about the trajectory rather than the quarter, and it is the same argument as OPS 1.4.
Verify. What proportion of your team’s capacity currently goes to toil, and what proportion is reserved for eliminating it? Which number is larger, and is that a decision?
Related
Section titled “Related”OPS 1.4Funding operational work, which this depends on entirelyOPS 3Everything as code, which is the largest single reduction in manual workOPS 8Operational procedures, the alternative when automation is not warrantedSEC 8Hardening and patching, which Server Update Management addressesSUS 6Shutting down idle, where the answer is deletion rather than automation