---
type: tradeoffs
pillar: operational-excellence
code: OPS
title: "Operational Excellence: tradeoffs"
description: What operational excellence costs up front, why its benefit arrives as an absence of incidents rather than a feature, and why both sides compound over time.
status: draft
sidebar:
  label: "Tradeoffs"
  order: 2
source_url: "https://framework.stackit.cloud/architecture/pillars/operational-excellence/tradeoffs/"
source_file: "docs/architecture/pillars/operational-excellence/tradeoffs.mdx"
---

Operational Excellence has an unusual cost profile. Most of its price is paid in engineering time
that produces nothing a customer can see, up front, before the benefit exists, and the benefit
arrives as an absence: incidents that did not happen, nights nobody was woken, migrations that
were uneventful.

Costs that are visible and immediate lose arguments to benefits that are invisible and deferred.
Which is why this pillar is the one most consistently under-funded, and why the under-funding is
usually described as pragmatism.

---

## Against Cost Optimization

**The most common conflict, and the one where "later" is most often the answer.**

Automation, pipelines, observability tooling, test infrastructure, and rehearsal all consume
engineering capacity that could have shipped features. Under delivery pressure they are deferred
first, and the deferral compounds: the team that had no time to automate now has less time,
because it is doing manually what the automation would have done.

Observability has its own cost shape. Telemetry volume grows with the system, its bill is
itemized, and its value only materializes during an incident. Retention gets cut in cost reviews
and the consequence appears during the next investigation, several months later, in a way nobody
attributes to the decision.

**How to resolve it:** count the labour. A cost model that omits who operates the thing (see
[[`COST 1`](/architecture/pillars/cost-optimization/cost-01-cost-model/)](/architecture/pillars/cost-optimization/)) makes managed services look expensive and manual
operations look free. For telemetry specifically, tier by signal value rather than cutting
uniformly: high-cardinality debugging data can have a short retention while the signals used for
trend analysis and audit are kept.

## Against Security

Two directions, and both are real.

**Operations wants access; security wants it constrained.** Broad access makes diagnosis fast.
Short-lived credentials, least privilege, and time-bound elevation all slow an engineer down at
the moment they are trying to fix something. This tension resolves badly by default: into a shared
account with standing permissions that everyone uses and nobody logs.

**Automation is privileged, and therefore a target.** A deployment pipeline can change production;
whoever controls it controls the system. Pipelines that were built for convenience frequently hold
broader permissions than any human, with weaker controls than any human account. This is one of
the more attractive paths into an environment.

**How to resolve it:** invest in fast, logged, legitimate paths, self-service time-bound
elevation, break-glass that is quick and reviewed afterwards: rather than relying on people to
tolerate friction. And treat the pipeline as production infrastructure with production-grade
access control, because it is.

Pulling the other way: [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) is a security asset. Infrastructure as code makes configuration
reviewable, auditable, and diffable, which is worth more than most dedicated security tooling.

## Against Performance Efficiency

Small and mostly mechanical. Instrumentation costs some CPU and some latency; tracing at full
sample rate on a hot path is measurable. Progressive deployment means running two versions
simultaneously, briefly, at some capacity cost.

**How to resolve it:** sample rather than remove. Adaptive sampling: full detail on errors and
slow requests, a fraction of the rest: retains almost all diagnostic value at a fraction of the
overhead.

## Against Reliability

Mostly aligned; two frictions worth naming.

**Change is the leading cause of incidents,** and this pillar advocates deploying more often. The
resolution is [`OPS 5`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/): the risk of frequent deployment comes from unmanaged exposure, not from
frequency. Small changes with progressive rollout and automatic rollback are safer per change and
per unit of time than large infrequent ones, but only if the safe-deployment machinery actually
exists. Frequent deployment without it is worse than infrequent deployment.

**Automation fails in ways humans do not.** An automated remediation with a bad condition applies
its mistake everywhere, immediately, at machine speed. Blast radius limits and circuit breakers
apply to automation as much as to application code.

## Against Sovereignty & Compliance

Three specific frictions.

**Key lifecycle is an operating commitment.** Customer-managed keys under [`SOV 4`](/architecture/pillars/sovereignty/sov-04-key-ownership/) mean operating
generation, rotation, revocation and backup of key material, with a procedure for the day a key
that a live database depends on is deleted. The
[Sovereignty tradeoffs](/architecture/pillars/sovereignty/tradeoffs/) call this the largest cost that pillar
imposes here, and the least anticipated.

**Telemetry is data.** Observability wants everything collected and correlated; [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/) requires
that logs and traces respect the classification of what they describe. Debug logs that dump
request bodies are the standard way regulated data ends up in an unclassified store.

**Restricted operator access cuts both ways.** Closing provider access paths for sovereignty
reasons removes support capabilities you may want during an incident. That is a decision to make
deliberately, in advance, rather than discover at the worst moment.

Pulling the other way, strongly: [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) and [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) produce most of what [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) requires.
Infrastructure as code proves what was deployed and when; correlated activity logs are the audit
trail. Compliance evidence as a by-product of good operations is largely this.

## Against Sustainability

Minor, and it runs both ways. Telemetry storage and multi-version deployments consume resources.
Against that, [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) and [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/) are what make automated shutdown of idle environments possible
at all, most of the Sustainability pillar depends on automation that this pillar builds.

---

## The compounding argument

Every tradeoff above is a comparison at a point in time, and that understates the case.

Operational Excellence compounds. Automation built once runs indefinitely. A rehearsed procedure
gets faster each time. Each incident review that finds a structural cause removes a class of
failure rather than an instance. Meanwhile its absence compounds in the other direction: manual
operations consume the time that would have automated them, undetected drift makes environments
diverge, and unreviewed incidents recur.

Which means the honest framing is not "this quarter's features versus this quarter's tooling". It
is a decision about the slope of the next two years. That does not always win the argument, and it
is the argument actually being had.

---

## Related

- <LinkChip href="/architecture/pillars/operational-excellence/principles/">Design principles</LinkChip>
- [Overview](/architecture/pillars/operational-excellence/): the questions this pillar asks
- <LinkChip href="/architecture/pillars/reliability/tradeoffs/">Reliability tradeoffs</LinkChip>
