---
id: OPS10
pillar: operational-excellence
title: OPS 10. How do you find and eliminate toil?
description: "Toil disguises itself as work: visible, appreciated, and consuming exactly the capacity that would have automated it away. How to measure and remove the need."
status: draft
services: [automation-service, compute-engine]
sidebar:
  order: 19
  label: Toil elimination
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/ops-10-toil-elimination/"
source_file: "docs/architecture/pillars/operational-excellence/ops-10-toil-elimination.mdx"
---

Toil is manual work that recurs, produces no lasting value, and scales with the size of the
system. Applying a patch by hand, clearing a stuck queue every Tuesday, provisioning an account on
request, copying figures into a spreadsheet before a meeting.

It is dangerous precisely because it looks like productivity. It is visible, people thank you for
it, and it feels like the job. Meanwhile it consumes the capacity that would have removed it,
which is why teams with the most toil have the least time to address it.

## Best practices

- [`OPS 10.1`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-101-measure-recurring-manual-work-rather-than-absorbing-it) Measure recurring manual work rather than absorbing it
- [`OPS 10.2`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-102-automate-what-recurs-starting-with-the-riskiest-rather-than-the-most-frequent) Automate what recurs, starting with the riskiest rather than the most frequent
- [`OPS 10.3`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-103-remove-the-need-rather-than-automating-the-workaround) Remove the need rather than automating the workaround
- [`OPS 10.4`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-104-budget-for-elimination-explicitly) Budget for elimination explicitly

---

## OPS 10.1 Measure recurring manual work rather than absorbing it

**Risk if not established:** Medium

Toil is invisible in every reporting system because nobody records it. It happens between the
tickets, and the team absorbs it until the absorption fails.

Measure it for a bounded period rather than forever. Two weeks of everyone noting recurring manual
operations and roughly how long each took produces a list that is uncomfortable and actionable.
Precision is not the point; the ranking is.

Capture three things per item: how often it recurs, how long it takes, and what happens if it is
done wrong. The third is what separates the annoying from the dangerous, and it is usually the one
nobody has considered.

Watch for the work that is invisible even to the people doing it: the mental overhead of
remembering that something must be checked, the interruptions, the context switching. Those cost
more than the minutes suggest and they are the reason a small recurring task can dominate a week.

**On STACKIT.** No platform feature measures your manual work.

Two sources give partial visibility. The
<LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> records actions taken through the
Portal, the CLI and the API, so a repeated manual action against platform resources is visible
there, within the retention window [`OPS 7.4`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-74-set-retention-by-the-value-of-each-signal-rather-than-uniformly) covers. And a high ratio of console actions to pipeline actions is itself
a signal, both of toil and of the drift [`OPS 3.3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-33-detect-drift-and-treat-it-as-a-defect) looks for.

**Tradeoffs.** **Operational Excellence.** Two weeks of light record-keeping. The main resistance is
that the exercise makes visible how much time goes to work nobody planned, which is the point and is
uncomfortable.

**Verify.** List the recurring manual operations your team performed last month, with frequency
and duration. What proportion of the team's time do they represent?

---

## OPS 10.2 Automate what recurs, starting with the riskiest rather than the most frequent

**Risk if not established:** Medium

The instinct is to automate the most frequent task, because the time saved is easiest to
calculate. The better first target is usually the one where a mistake is most expensive.

A weekly task that saves twenty minutes is worth automating. A quarterly task that touches
production data, has eleven steps, and has gone wrong twice is worth automating first, because the
value is in eliminating the error rather than the minutes.

Automation should fail safely and be observable. A script that stops and alerts is better than one
that continues past an unexpected state, and one that reports what it did is what makes it
trustworthy enough to leave alone. Automation with a bad condition applies its mistake everywhere
at machine speed, which is the point [`OPS 5.3`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-53-roll-back-automatically-when-a-gate-fails) also makes.

Not everything should be automated. Work that is genuinely different each time, needs judgement,
or happens twice a year and takes ten minutes is better documented than automated, since the
automation would itself become a thing to maintain. [`OPS 8`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/) is the alternative, and choosing
between the two deliberately is part of this.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/integration/automation-service/">Automation
Service</LinkChip> is the managed
platform for scheduled and triggered automation across STACKIT services, with
<LinkChip href="https://docs.stackit.cloud/products/integration/automation-service/how-tos/automation-templates/">templates</LinkChip>
for recurring patterns.

For virtual machines, <LinkChip href="https://docs.stackit.cloud/products/compute-engine/run-command/">Run
Command</LinkChip> executes scripts across
servers with <LinkChip href="https://docs.stackit.cloud/products/compute-engine/run-command/getting-started/command-templates/">command
templates</LinkChip>,
which covers a large share of the repetitive operational work on Compute Engine.

Patching is one category worth naming, because it is high-frequency and high-consequence. <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-update-management/">Server
Update Management</LinkChip>
provides scheduled operating system updates for Linux and Windows, which removes a recurring
manual operation and satisfies [`SEC 8`](/architecture/pillars/security/sec-08-hardening-and-patching/) at the same time.

Beyond those, the <LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-cli/">CLI</LinkChip>,
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-api/">API</LinkChip> and
<LinkChip href="https://docs.stackit.cloud/developer-tools/stackit-sdk/stackit-go-sdk/">SDKs</LinkChip> mean anything
reachable through the platform can be scripted, which is the same reach [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) depends on.

**Tradeoffs.** **Operational Excellence**, against itself: automation is code that must be
maintained, tested and understood. Automating something that changes frequently produces a
maintenance burden larger than the toil it replaced.

**Verify.** Of the recurring operations you measured, which have been automated in the last six
months? Which was chosen first, and was the reason frequency or consequence?

---

## OPS 10.3 Remove the need rather than automating the workaround

**Risk if not established:** Medium

Automating a recurring task makes it cheaper. Removing the reason it recurs makes it free, and
the second is available more often than teams assume.

A queue that has to be cleared weekly is telling you about a design problem. A service that must
be restarted every few days has a leak. A certificate that needs manual renewal has an automation
gap upstream. Automating any of those hides the signal and makes the underlying defect permanent,
because nobody feels it any more.

Ask why before asking how. If a task recurs because of a design decision, changing the design
removes it entirely. If it recurs because of a missing capability, building that capability
removes a class of tasks rather than one.

The pattern to look for is toil that keeps reappearing in the same component. [`OPS 3.4`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-34-codify-emergency-changes-afterwards) makes the
same observation about emergency changes: a component that repeatedly needs manual intervention is
describing itself.

Sometimes the answer is to delete the thing. A report nobody reads, an environment nobody opens, a
job whose output goes nowhere. That is [`SUS 6`](/architecture/pillars/sustainability/sus-06-shut-down-idle/) and it is the cheapest fix available.

**On STACKIT.** No platform feature does this. It is a design decision each time.

The platform-shaped version of the question is worth asking though: some toil exists because
something is self-operated that could be consumed as a managed service. Moving a self-run database
to <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/">PostgreSQL Flex</LinkChip> or a self-run
CI system to <LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/basics/stackit-pipelines/">STACKIT
Pipelines</LinkChip>
removes the operational work rather than automating it. The comparison belongs in [`COST 1`](/architecture/pillars/cost-optimization/cost-01-cost-model/), which
routinely omits the labour on the self-operated side.

**Tradeoffs.** **Cost Optimization.** Removing a need is usually a larger change than automating a
task, and it competes with feature work on a longer timescale. It is also the only option that
scales.

**Verify.** For your three largest sources of toil, why does each recur? For how many is the
answer a design decision that could be changed rather than a task that could be scripted?

---

## OPS 10.4 Budget for elimination explicitly

**Risk if not established:** Medium

Toil elimination has the worst possible shape for getting funded: the cost is immediate and
visible, the benefit is deferred and diffuse, and the work produces nothing a customer sees.

Left to compete with features on merit, it loses every time, and the loss compounds. Each quarter
of unaddressed toil consumes more of the following quarter, which is the spiral described in
[`OPS 1.4`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-14-fund-operational-work-explicitly-rather-than-expecting-it-to-fit-in-the-gaps).

Reserve capacity rather than intending to find it. A stated proportion, protected in the same way
feature commitments are protected, is the only version that survives a deadline. The number
matters less than the fact that it is a decision.

Set a ceiling as a trigger. Where toil exceeds an agreed share of capacity, elimination takes
priority over new work until it is back under. That converts a judgement call into a rule, which
is what makes it survive the quarter where everything is urgent.

Report the result. Toil eliminated is capacity returned, and it is the only form in which this
work shows up as a number anyone outside the team recognizes.

**On STACKIT.** No platform feature applies.

The one platform decision that changes the arithmetic is managed versus self-operated, since it
moves operational work off your team permanently. That comparison is only honest if the labour is
counted, which is [`COST 1`](/architecture/pillars/cost-optimization/cost-01-cost-model/) and the reason a managed database that looks expensive next to a
virtual machine is frequently cheaper.

**Tradeoffs.** **Cost Optimization.** Reserved capacity is capacity not spent on features. The
argument is about the trajectory rather than the quarter, and it is the same argument as `OPS
1.4`.

**Verify.** What proportion of your team's capacity currently goes to toil, and what proportion is
reserved for eliminating it? Which number is larger, and is that a decision?

---

## Related

- [`OPS 1.4`](/architecture/pillars/operational-excellence/ops-01-shared-ownership/#ops-14-fund-operational-work-explicitly-rather-than-expecting-it-to-fit-in-the-gaps) Funding operational work, which this depends on entirely
- [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) Everything as code, which is the largest single reduction in manual work
- [`OPS 8`](/architecture/pillars/operational-excellence/ops-08-operational-procedures/) Operational procedures, the alternative when automation is not warranted
- [`SEC 8`](/architecture/pillars/security/sec-08-hardening-and-patching/) Hardening and patching, which Server Update Management addresses
- [`SUS 6`](/architecture/pillars/sustainability/sus-06-shut-down-idle/) Shutting down idle, where the answer is deletion rather than automation
