Zum Inhalt springen
Beta

Performance Efficiency: design principles

Zuletzt aktualisiert am

1. Performance targets come from users, not from benchmarks

Section titled “1. Performance targets come from users, not from benchmarks”

“Fast” is not a specification. Without a stated target, performance work has no completion criterion and no failure criterion, it continues until someone loses interest or a budget runs out, and nobody can say whether the result is adequate.

The target comes from the people the system serves: what the interaction actually needs to feel responsive, what a downstream system requires to meet its own commitment, what a batch process must finish before to be useful. These are business inputs. They are not derivable from infrastructure capability, and a benchmark result tells you what the system does, never what it should do.

Targets are also per flow, not per workload. The checkout path and the monthly report have legitimately different requirements, and a single number applied to both is wrong for at least one.

2. Measure, then change: never the reverse

Section titled “2. Measure, then change: never the reverse”

Intuition about performance is unreliable in a specific, well-documented way: the bottleneck is rarely where the design discussion assumed. Systems have too many interacting parts, caches behave unexpectedly, and the code that looks expensive is often called once while the trivial function is called a million times.

An unmeasured optimization is a guess that costs engineering time, adds permanent complexity, and delivers an unverified benefit. Sometimes zero benefit. Occasionally negative, an added cache that introduced a consistency bug and was protecting something that had not been slow since a schema change two years earlier.

Measure before, to find the constraint. Measure after, to confirm it moved. If it did not, revert: complexity without benefit is pure cost, and it is much easier to remove now than in a year.

When something is slow, the available responses are to make it do less work or to give it more resources. The second is faster to implement and is usually reached for first.

It also has properties worth noticing. It costs money permanently, scales the problem rather than solving it, and hides the underlying inefficiency so that the next enlargement helps less than the last. The classic sequence is a database that gets progressively larger instances until someone finally examines the query and discovers it was reading the entire table.

Work not done is free, needs no maintenance, consumes no capacity, and never regresses. Look for it first, the unused columns in the query, the payload nobody parses, the recomputation of something that has not changed, the round trip that could have been one call. Then scale, having established that the remaining work genuinely requires it.

Systems have a bottleneck: usually one, at any given moment. Adding capacity anywhere else changes nothing except the bill.

This gets missed because scaling is often applied at the layer that is easiest to scale rather than the one that is saturated. Application servers scale readily; the database behind them does not. Doubling the application tier when the constraint is a database connection pool produces more contention and no throughput.

Find the constraint, relieve it, then find the next one, because relieving a bottleneck always reveals another. That is not failure; it is what progress looks like. The mistake is assuming the first constraint was the only one and stopping the measurement.

A system that met its targets at launch will not meet them indefinitely, and no single change will be responsible.

Data volumes grow, and query plans that were correct at ten thousand rows are wrong at ten million. Features accumulate, each adding a small amount of work to a common path. Dependencies update. Traffic patterns shift as usage matures. Every increment is individually negligible; together they are the reason systems get slower without anyone breaking anything.

Which makes performance a continuous property, not a launch-time achievement. It needs a baseline (PERF 2), automated detection of regression, and periodic re-examination of decisions whose justification may have expired. Some optimizations should be removed over time, the cache that no longer protects anything, the denormalization that a schema change made pointless. Complexity retained past its usefulness is a cost with no remaining benefit.