Skip to content
Beta

Operational Excellence: tradeoffs

Last updated on

Operational Excellence has an unusual cost profile. Most of its price is paid in engineering time that produces nothing a customer can see, up front, before the benefit exists, and the benefit arrives as an absence: incidents that did not happen, nights nobody was woken, migrations that were uneventful.

Costs that are visible and immediate lose arguments to benefits that are invisible and deferred. Which is why this pillar is the one most consistently under-funded, and why the under-funding is usually described as pragmatism.


The most common conflict, and the one where “later” is most often the answer.

Automation, pipelines, observability tooling, test infrastructure, and rehearsal all consume engineering capacity that could have shipped features. Under delivery pressure they are deferred first, and the deferral compounds: the team that had no time to automate now has less time, because it is doing manually what the automation would have done.

Observability has its own cost shape. Telemetry volume grows with the system, its bill is itemized, and its value only materializes during an incident. Retention gets cut in cost reviews and the consequence appears during the next investigation, several months later, in a way nobody attributes to the decision.

How to resolve it: count the labour. A cost model that omits who operates the thing (see [COST 1](/architecture/pillars/cost-optimization/)) makes managed services look expensive and manual operations look free. For telemetry specifically, tier by signal value rather than cutting uniformly: high-cardinality debugging data can have a short retention while the signals used for trend analysis and audit are kept.

Two directions, and both are real.

Operations wants access; security wants it constrained. Broad access makes diagnosis fast. Short-lived credentials, least privilege, and time-bound elevation all slow an engineer down at the moment they are trying to fix something. This tension resolves badly by default: into a shared account with standing permissions that everyone uses and nobody logs.

Automation is privileged, and therefore a target. A deployment pipeline can change production; whoever controls it controls the system. Pipelines that were built for convenience frequently hold broader permissions than any human, with weaker controls than any human account. This is one of the more attractive paths into an environment.

How to resolve it: invest in fast, logged, legitimate paths, self-service time-bound elevation, break-glass that is quick and reviewed afterwards: rather than relying on people to tolerate friction. And treat the pipeline as production infrastructure with production-grade access control, because it is.

Pulling the other way: OPS 3 is a security asset. Infrastructure as code makes configuration reviewable, auditable, and diffable, which is worth more than most dedicated security tooling.

Small and mostly mechanical. Instrumentation costs some CPU and some latency; tracing at full sample rate on a hot path is measurable. Progressive deployment means running two versions simultaneously, briefly, at some capacity cost.

How to resolve it: sample rather than remove. Adaptive sampling: full detail on errors and slow requests, a fraction of the rest: retains almost all diagnostic value at a fraction of the overhead.

Mostly aligned; two frictions worth naming.

Change is the leading cause of incidents, and this pillar advocates deploying more often. The resolution is OPS 5: the risk of frequent deployment comes from unmanaged exposure, not from frequency. Small changes with progressive rollout and automatic rollback are safer per change and per unit of time than large infrequent ones, but only if the safe-deployment machinery actually exists. Frequent deployment without it is worse than infrequent deployment.

Automation fails in ways humans do not. An automated remediation with a bad condition applies its mistake everywhere, immediately, at machine speed. Blast radius limits and circuit breakers apply to automation as much as to application code.

Three specific frictions.

Key lifecycle is an operating commitment. Customer-managed keys under SOV 4 mean operating generation, rotation, revocation and backup of key material, with a procedure for the day a key that a live database depends on is deleted. The Sovereignty tradeoffs call this the largest cost that pillar imposes here, and the least anticipated.

Telemetry is data. Observability wants everything collected and correlated; SOV 3 requires that logs and traces respect the classification of what they describe. Debug logs that dump request bodies are the standard way regulated data ends up in an unclassified store.

Restricted operator access cuts both ways. Closing provider access paths for sovereignty reasons removes support capabilities you may want during an incident. That is a decision to make deliberately, in advance, rather than discover at the worst moment.

Pulling the other way, strongly: OPS 3 and OPS 7 produce most of what SOV 7 requires. Infrastructure as code proves what was deployed and when; correlated activity logs are the audit trail. Compliance evidence as a by-product of good operations is largely this.

Minor, and it runs both ways. Telemetry storage and multi-version deployments consume resources. Against that, OPS 3 and OPS 10 are what make automated shutdown of idle environments possible at all, most of the Sustainability pillar depends on automation that this pillar builds.


Every tradeoff above is a comparison at a point in time, and that understates the case.

Operational Excellence compounds. Automation built once runs indefinitely. A rehearsed procedure gets faster each time. Each incident review that finds a structural cause removes a class of failure rather than an instance. Meanwhile its absence compounds in the other direction: manual operations consume the time that would have automated them, undetected drift makes environments diverge, and unreviewed incidents recur.

Which means the honest framing is not “this quarter’s features versus this quarter’s tooling”. It is a decision about the slope of the next two years. That does not always win the argument, and it is the argument actually being had.