Performance Efficiency
Zuletzt aktualisiert am
Does it meet demand without waste?
Both halves of that question matter, and the second is what makes this a distinct pillar. Meeting demand is easy if you are permitted to provision without limit. Performance efficiency is meeting demand while the resources consumed stay proportionate to the work delivered, which is why the pillar sits between raw capacity and cost, borrowing from both and reducible to neither.
The discipline is unusual in one respect: it is the pillar where intuition performs worst. Engineers are systematically bad at predicting where time is spent in their own systems. The bottleneck is almost never where the design discussion assumed, optimizations are frequently applied to code that was not the constraint, and the fix that would have mattered was in a place nobody was looking. This is not a failure of skill. It is what happens when a system has more interacting parts than anyone can hold in mind, and it is why measurement is not a step in this pillar but its foundation.
What this pillar covers
Section titled “What this pillar covers”- Deriving performance targets from user expectations
- Establishing and maintaining a baseline
- Selecting and sizing services from measured demand
- Designing data access for the query pattern
- Scaling: horizontally where possible, on signals that predict rather than react
- Reducing work, caching, batching, compression, locality
- Optimizing on profiling evidence
- Testing under realistic load, before and after production
- Maintaining performance over time as the system and its traffic change
What it does not cover
Section titled “What it does not cover”Whether the system stays available under failure belongs to Reliability, though the pillars meet where saturation becomes an outage. Whether the resources are worth their price belongs to Cost Optimization: the two mostly agree, since work not done is neither slow nor expensive. Whether the consumption is environmentally justified belongs to Sustainability.
The telemetry these questions are answered from belongs to Operational
Excellence, as does the deployment gate that catches a regression
before it reaches everyone. This pillar decides what the numbers should be, and what to change
when they move. Keeping performance good as the workload changes is therefore here, in PERF 9,
and not an operational practice.
The central idea
Section titled “The central idea”Measure, then change. Never the reverse.
An optimization applied without measurement is a guess with a cost: it consumes engineering time, adds complexity that has to be maintained forever, and produces a benefit that nobody verified. Frequently it produces no benefit at all, and occasionally a negative one, and either way the complexity stays. Every experienced engineer has removed an elaborate cache that was protecting something which had not been slow for two years.
The second idea: efficiency beats capacity. The instinct when a system is slow is to give it more, a bigger instance, more replicas, another cache. Sometimes correct. But the work not done is faster and cheaper than the work done efficiently on more hardware, and it does not need to be maintained. A query that stops requesting data it never uses outperforms every amount of hardware thrown at the version that did.
Scaling hardware also hides the problem rather than removing it, which is why the third enlargement usually helps less than the first.
Where to start
Section titled “Where to start”Questions
Section titled “Questions”Nine questions. Numbers follow the order the decisions are usually made in and do not indicate priority.
PERF 1: How do you derive performance targets from user expectations?
Section titled “PERF 1: How do you derive performance targets from user expectations?”State the target per flow, at a defined percentile, under a defined load, and have it agreed by someone who represents the users. Without a target, performance work has no completion criterion and no failure criterion.
PERF 2: How do you establish a performance baseline and detect regressions?
Section titled “PERF 2: How do you establish a performance baseline and detect regressions?”Record how the system currently behaves under known conditions, and compare automatically as it changes. Performance decays through accumulated small increments that are invisible individually and obvious only against a baseline.
PERF 3: How do you select and size services from measured demand?
Section titled “PERF 3: How do you select and size services from measured demand?”Choose service types and sizes from observed load and its shape, not from a starting guess that was never revisited. Know where the scaling limit of each choice is before you approach it.
PERF 4: How do you design data access for the actual query pattern?
Section titled “PERF 4: How do you design data access for the actual query pattern?”Model, index, and partition for the queries the workload actually issues. Data access is the most common location of the real constraint and the place where architectural mistakes are most expensive to correct later.
PERF 5: How do you scale, and on which signals?
Section titled “PERF 5: How do you scale, and on which signals?”Prefer adding instances to enlarging them, which requires designing components to be stateless or to partition cleanly. Trigger scaling on leading indicators such as queue depth rather than on lagging ones such as CPU, and account for the time scaling takes.
PERF 6: How do you reduce the work the system does?
Section titled “PERF 6: How do you reduce the work the system does?”The cheapest work is the work not done: cache what is expensive and stable, batch what is chatty, compress what travels, and move computation to the data rather than the reverse. Each of these adds complexity, so each needs a measured justification.
PERF 7: How do you decide what to optimize?
Section titled “PERF 7: How do you decide what to optimize?”Profile to find the constraint, change one thing, measure whether the constraint moved. If it did not, revert, complexity without benefit is pure cost and is far easier to remove now than in a year.
PERF 8: How do you test performance under realistic load?
Section titled “PERF 8: How do you test performance under realistic load?”Load-test with realistic data volumes, realistic distributions, and realistic concurrency, and include the failure region, how the system behaves past saturation is a design property, not an accident. Keep testing after launch, since production conditions drift away from any single test.
PERF 9: How do you maintain performance over the workload’s life?
Section titled “PERF 9: How do you maintain performance over the workload’s life?”Re-examine sizing, indexes, and caching on a cadence, and remove optimizations whose justification has expired. Retained complexity that no longer buys anything is a cost with no remaining benefit.
Related
Section titled “Related”- Design principles
- Tradeoffs
- Reliability:
PERF 5andREL 7are the same mechanism serving different objectives