PERF 2. How do you establish a performance baseline and detect regressions?
Zuletzt aktualisiert am
Systems get slower without anyone breaking anything. Data grows, features accumulate on a common path, dependencies update, traffic patterns shift. Each increment is negligible and the sum is not.
A baseline is what converts that from a vague sense that things feel slower into a measurement. Without one, the first reliable signal is a complaint, by which point the cause is months of changes rather than one.
Best practices
Section titled “Best practices”PERF 2.1Record how the system behaves under known conditionsPERF 2.2Measure percentiles rather than averagesPERF 2.3Compare automatically rather than by periodic reviewPERF 2.4Keep the baseline current as the system changes
PERF 2.1 Record how the system behaves under known conditions
Section titled “PERF 2.1 Record how the system behaves under known conditions”Risk if not established: Medium
A measurement without its conditions cannot be compared with anything. The baseline needs the conditions recorded alongside the numbers: what load, what data volume, which version, which environment, and what else was running.
Baseline the flows from REL 2.2 rather than every endpoint. A baseline covering everything is
expensive to maintain and nobody reads it; one covering the ranked flows gets looked at.
Capture more than latency. Throughput at that latency, error rate, and the resource utilization that produced it. A flow meeting its latency target at ninety percent CPU is one change away from missing it, and the latency alone does not say that.
Establish it early. A baseline taken after a system has been in production for two years cannot tell you what it used to do, which is exactly the question that matters when someone asks why it feels slow.
On STACKIT. Observability stores the measurements, and its service plans determine how far back a comparison can reach: metrics default to 90 days and can be extended to 26 months. For trend analysis over a system’s life the longer setting is the one that matters, and it is a decision to make before the history you want is gone.
Managed services publish which metrics they expose, for example for PostgreSQL
Flex ,
which is worth reading before designing a baseline around a metric that does not exist. What the
platform emits about itself and what your application emits under OPS 7 are separate sources and
both belong in the baseline.
Tradeoffs. Cost Optimization. Longer metric retention costs more, and the value only appears when someone asks a question about the past. Retention is the cheapest part of this question and the easiest to cut.
Verify. For your most critical flow, what was its 95th percentile latency six months ago? If you cannot answer, what would you compare today’s measurement against?
PERF 2.2 Measure percentiles rather than averages
Section titled “PERF 2.2 Measure percentiles rather than averages”Risk if not established: Medium
An average is dominated by the common case and says nothing about the tail. A system averaging 200 milliseconds might have a well-behaved distribution or might serve one request in twenty in four seconds, and those are entirely different systems to the people using them.
Track at least the median, a high percentile such as the 95th, and an extreme such as the 99th. The median tells you about the typical experience, the high percentile about the experience people complain about, and the gap between them tells you whether the system is consistent.
Averages also hide the shape of a regression. A change that makes ten percent of requests much slower barely moves the average and moves the 95th percentile sharply, which is the difference between noticing and not.
Watch the tail specifically for the things that cause it: garbage collection, cache misses, lock
contention, retries under REL 5.2, and cold starts. These are invisible in the mean and are most
of what users experience as unreliability.
On STACKIT. Percentile calculation is a property of your metric instrumentation rather than of
the platform. What you emit under OPS 7.1 determines whether percentiles are available at all: a
metric recorded as a single average value cannot be decomposed afterwards.
Tradeoffs. Cost Optimization. Percentile metrics carry more data than a single average,
and high-cardinality dimensions multiply that, which is the constraint OPS 7.3 describes.
Verify. For your critical flows, which percentiles are recorded? What is the ratio between the median and the 95th, and has that ratio changed?
PERF 2.3 Compare automatically rather than by periodic review
Section titled “PERF 2.3 Compare automatically rather than by periodic review”Risk if not established: Medium
A baseline reviewed quarterly detects a regression up to three months after it arrived, by which point the change that caused it is buried among hundreds of others.
Automate the comparison at the two points where it is cheapest. Before merge, where a performance test in the pipeline can catch an obvious regression against a known workload, and where the suspect is one change. In production, comparing against the recent trend, which catches what only appears under real load and real data.
The second is where most regressions are actually found, because the conditions that produce them
rarely exist in a test. This is the same argument SEC 1.2 makes about pre-deployment and
post-deployment measurement.
Set the threshold from consequence rather than from a round percentage. A flow far from its target can absorb a ten percent regression; one already close to its limit cannot absorb three.
Expect noise. Performance measurements vary between runs, and a comparison without a tolerance
produces alerts nobody trusts, which is the fatigue problem from REL 10.2 in a different form.
On STACKIT. Alerting in
Observability
is where the production-side comparison is expressed. The pipeline-side check runs in STACKIT
Pipelines ,
using whichever load tool you choose, and it is the same mechanism as the health gate in OPS 5.2.
For the pipeline result to mean anything, the environment it runs in has to resemble production in
the ways that affect performance, which is OPS 6 and specifically the resource shape from
PERF 3.
Tradeoffs. Operational Excellence. Performance tests in a pipeline are slow, and slow
pipelines push teams toward batching changes, which works against OPS 4.3. Run the full
comparison on a schedule and a fast subset per change.
Verify. How would you find out that your critical flow became 30% slower this week? How long would that take, and what would tell you?
PERF 2.4 Keep the baseline current as the system changes
Section titled “PERF 2.4 Keep the baseline current as the system changes”Risk if not established: Medium
A baseline that is never updated becomes a historical curiosity. The system it described no longer exists, and comparing against it produces differences that reflect intended changes rather than regressions.
Update it deliberately when something legitimately changes the expected behaviour: a new feature on the path, a deliberate capacity change, a platform version. Record why, so that a future reader can tell an intended shift from an accepted degradation.
The discipline that matters is separating the two. A regression that gets absorbed into the baseline because nobody questioned it is a regression that has become permanent, and the baseline now certifies it. Requiring a reason for every baseline change is what prevents that.
Keep the old values. The interesting question is frequently how the system behaved over years rather than against the last revision, and that history is only available if it was not overwritten.
On STACKIT. The metrics history in
Observability
is what makes the long comparison possible, subject to the retention decided in PERF 2.1. Where
the baseline itself is a recorded artefact rather than a query, it belongs in version control
under OPS 3.1, so that changes to it are reviewed like any other.
Tradeoffs. Operational Excellence. Another artefact with an owner and a review, and one that is easy to let drift because nothing breaks when it does.
Verify. When was your baseline last updated, and what changed to justify it? Was any of that change a regression that got accepted rather than fixed?
Related
Section titled “Related”PERF 1Targets, which the baseline is compared againstPERF 7Evidence-based optimization, which starts from the baselinePERF 9Performance lifecycle, where accumulated drift is addressedOPS 7Observability, which supplies the measurementsOPS 5.2Health gates, which use the same comparison mechanism