Operational Excellence: tradeoffs
Last updated on
Operational Excellence has an unusual cost profile. Most of its price is paid in engineering time that produces nothing a customer can see, up front, before the benefit exists, and the benefit arrives as an absence: incidents that did not happen, nights nobody was woken, migrations that were uneventful.
Costs that are visible and immediate lose arguments to benefits that are invisible and deferred. Which is why this pillar is the one most consistently under-funded, and why the under-funding is usually described as pragmatism.
Against Cost Optimization
Section titled “Against Cost Optimization”The most common conflict, and the one where “later” is most often the answer.
Automation, pipelines, observability tooling, test infrastructure, and rehearsal all consume engineering capacity that could have shipped features. Under delivery pressure they are deferred first, and the deferral compounds: the team that had no time to automate now has less time, because it is doing manually what the automation would have done.
Observability has its own cost shape. Telemetry volume grows with the system, its bill is itemized, and its value only materializes during an incident. Retention gets cut in cost reviews and the consequence appears during the next investigation, several months later, in a way nobody attributes to the decision.
How to resolve it: count the labour. A cost model that omits who operates the thing (see
[COST 1](/architecture/pillars/cost-optimization/)) makes managed services look expensive and manual
operations look free. For telemetry specifically, tier by signal value rather than cutting
uniformly: high-cardinality debugging data can have a short retention while the signals used for
trend analysis and audit are kept.
Against Security
Section titled “Against Security”Two directions, and both are real.
Operations wants access; security wants it constrained. Broad access makes diagnosis fast. Short-lived credentials, least privilege, and time-bound elevation all slow an engineer down at the moment they are trying to fix something. This tension resolves badly by default: into a shared account with standing permissions that everyone uses and nobody logs.
Automation is privileged, and therefore a target. A deployment pipeline can change production; whoever controls it controls the system. Pipelines that were built for convenience frequently hold broader permissions than any human, with weaker controls than any human account. This is one of the more attractive paths into an environment.
How to resolve it: invest in fast, logged, legitimate paths, self-service time-bound elevation, break-glass that is quick and reviewed afterwards: rather than relying on people to tolerate friction. And treat the pipeline as production infrastructure with production-grade access control, because it is.
Pulling the other way: OPS 3 is a security asset. Infrastructure as code makes configuration
reviewable, auditable, and diffable, which is worth more than most dedicated security tooling.
Against Performance Efficiency
Section titled “Against Performance Efficiency”Small and mostly mechanical. Instrumentation costs some CPU and some latency; tracing at full sample rate on a hot path is measurable. Progressive deployment means running two versions simultaneously, briefly, at some capacity cost.
How to resolve it: sample rather than remove. Adaptive sampling: full detail on errors and slow requests, a fraction of the rest: retains almost all diagnostic value at a fraction of the overhead.
Against Reliability
Section titled “Against Reliability”Mostly aligned; two frictions worth naming.
Change is the leading cause of incidents, and this pillar advocates deploying more often. The
resolution is OPS 5: the risk of frequent deployment comes from unmanaged exposure, not from
frequency. Small changes with progressive rollout and automatic rollback are safer per change and
per unit of time than large infrequent ones, but only if the safe-deployment machinery actually
exists. Frequent deployment without it is worse than infrequent deployment.
Automation fails in ways humans do not. An automated remediation with a bad condition applies its mistake everywhere, immediately, at machine speed. Blast radius limits and circuit breakers apply to automation as much as to application code.
Against Sovereignty & Compliance
Section titled “Against Sovereignty & Compliance”Three specific frictions.
Key lifecycle is an operating commitment. Customer-managed keys under SOV 4 mean operating
generation, rotation, revocation and backup of key material, with a procedure for the day a key
that a live database depends on is deleted. The
Sovereignty tradeoffs call this the largest cost that pillar
imposes here, and the least anticipated.
Telemetry is data. Observability wants everything collected and correlated; SOV 3 requires
that logs and traces respect the classification of what they describe. Debug logs that dump
request bodies are the standard way regulated data ends up in an unclassified store.
Restricted operator access cuts both ways. Closing provider access paths for sovereignty reasons removes support capabilities you may want during an incident. That is a decision to make deliberately, in advance, rather than discover at the worst moment.
Pulling the other way, strongly: OPS 3 and OPS 7 produce most of what SOV 7 requires.
Infrastructure as code proves what was deployed and when; correlated activity logs are the audit
trail. Compliance evidence as a by-product of good operations is largely this.
Against Sustainability
Section titled “Against Sustainability”Minor, and it runs both ways. Telemetry storage and multi-version deployments consume resources.
Against that, OPS 3 and OPS 10 are what make automated shutdown of idle environments possible
at all, most of the Sustainability pillar depends on automation that this pillar builds.
The compounding argument
Section titled “The compounding argument”Every tradeoff above is a comparison at a point in time, and that understates the case.
Operational Excellence compounds. Automation built once runs indefinitely. A rehearsed procedure gets faster each time. Each incident review that finds a structural cause removes a class of failure rather than an instance. Meanwhile its absence compounds in the other direction: manual operations consume the time that would have automated them, undetected drift makes environments diverge, and unreviewed incidents recur.
Which means the honest framing is not “this quarter’s features versus this quarter’s tooling”. It is a decision about the slope of the next two years. That does not always win the argument, and it is the argument actually being had.
Related
Section titled “Related”- Design principles
- Overview: the questions this pillar asks
- Reliability tradeoffs