Zum Inhalt springen
Beta

Performance Efficiency: tradeoffs

Zuletzt aktualisiert am

Performance has a particular way of losing arguments and then winning them badly. It is deferred during design because nothing is slow yet, deferred during delivery because features come first, and then addressed urgently under production pressure, at which point the cheap architectural options are gone and what remains is hardware and complexity.

The costs below are worth paying where a stated target requires them. Applied without a target, they are complexity purchased for no agreed benefit, and PERF 1 exists precisely to prevent that.


Mostly allies, see the Cost Optimization tradeoffs for the same ground from the other side. Work not done is neither slow nor expensive, and most genuine performance improvements reduce cost.

They part company in three places:

  • Headroom. Capacity that absorbs spikes is idle most of the time, and idle capacity is the first thing a cost review finds.
  • Caching and pre-computation buy latency with memory and storage. Whether the trade is good depends on hit rates and data volumes, both of which move.
  • Committed capacity is cheaper per unit and less elastic. A cost win that becomes a performance constraint the moment demand shifts.

How to resolve it: the stated target from PERF 1 does the same work REL 1 does for reliability. It converts “is this fast enough” from an opinion into a comparison. Without it, performance gets whatever the current cost pressure leaves over.

The cost that gets underestimated most. Nearly every performance mechanism adds a component, a failure mode, or a piece of state that has to be reasoned about.

Caches introduce invalidation, staleness, and an entire class of bug where the system is correct and the cache is not. Read replicas introduce replication lag, which application code must now handle. Sharding introduces rebalancing, cross-shard queries, and hot-partition management. Autoscaling introduces oscillation, cold starts, and thresholds that need tuning. Denormalization introduces multiple copies of a fact that can disagree.

Each of these is justified when a target requires it. Each is a permanent operational burden and a new way for the system to be subtly wrong at three in the morning.

How to resolve it: PERF 7’s discipline of reverting unproven optimizations is the main defence, and PERF 9 is the other half: periodically remove the mechanisms whose justification has expired. Complexity is easy to add incrementally and hard to remove later, so the removal has to be scheduled rather than hoped for.

Aligned in mechanism, divergent in intent, and the divergence is worth naming because the same tools serve both.

  • Scaling serves availability and throughput at once, and PERF 5 and REL 7 are largely the same question from two angles.
  • Synchronous replication costs write latency to buy durability. Asynchronous replication returns the latency and reopens the data-loss window. This is an RPO decision that frequently gets made as a performance decision.
  • Caching improves latency and can serve stale data during a failure, which is graceful degradation when deliberate and a correctness bug when accidental.
  • Aggressive timeouts protect the caller and abandon work the callee may have completed.
  • Retries improve success rates and add load precisely when the system is already struggling.

How to resolve it: decide replication mode per data set from the RPO, not from a latency benchmark. And be explicit about whether stale-cache behaviour under failure is a designed degradation path or an accident: the mechanism is identical and only one of them is safe.

Small and routinely overestimated in design discussions.

Encryption in transit costs handshake latency and some throughput, negligible for most workloads on modern hardware. Per-request authorization adds a lookup, addressed by caching, which lengthens the window in which a revoked permission still works. Inspection points add hops. Input validation costs cycles and is never worth removing.

How to resolve it: profile before assuming. Security overhead appears in performance arguments far more often than it appears in profiles. Where it is real, the answer is normally caching or a design change, not weakening the control.

  • Placement constraints limit how close compute can be to users. For a European user base this rarely binds; for a global one it does.
  • Customer-managed keys put a key service on paths that would otherwise be local: usually amortized by caching, occasionally material at high volume.
  • Confidential computing carries measurable overhead that varies by workload shape.
  • Data minimization sometimes removes exactly the data an optimization relied on.

How to resolve it: measure the specific case rather than reasoning from general figures, which vary too much to be useful. And treat the sovereignty requirement as the constraint the design works within, not as a variable to trade against latency.

The pillar with the most interesting relationship, because it is not simply alignment.

Efficiency serves both: less work means fewer resources and lower latency. PERF 6 and most of SUS 7 are the same question.

But performance is frequently bought with resources, and that is where they diverge. Headroom for spikes is idle capacity. Pre-computation trades storage for speed. Aggressive replication for read performance duplicates data. Keeping instances warm to avoid cold starts consumes resources continuously to serve occasional requests. Each is a legitimate performance decision and each consumes more to deliver the same work faster.

How to resolve it: be honest about which target requires the resource. Latency headroom for a flow with a stated target is justified; the same headroom on a batch process nobody is waiting for is waste that happens to look like engineering.