Skip to content
Beta

PERF 2. How do you establish a performance baseline and detect regressions?

Last updated on

Systems get slower without anyone breaking anything. Data grows, features accumulate on a common path, dependencies update, traffic patterns shift. Each increment is negligible and the sum is not.

A baseline is what converts that from a vague sense that things feel slower into a measurement. Without one, the first reliable signal is a complaint, by which point the cause is months of changes rather than one.

  • PERF 2.1 Record how the system behaves under known conditions
  • PERF 2.2 Measure percentiles rather than averages
  • PERF 2.3 Compare automatically rather than by periodic review
  • PERF 2.4 Keep the baseline current as the system changes

PERF 2.1 Record how the system behaves under known conditions

Section titled “PERF 2.1 Record how the system behaves under known conditions”

Risk if not established: Medium

A measurement without its conditions cannot be compared with anything. The baseline needs the conditions recorded alongside the numbers: what load, what data volume, which version, which environment, and what else was running.

Baseline the flows from REL 2.2 rather than every endpoint. A baseline covering everything is expensive to maintain and nobody reads it; one covering the ranked flows gets looked at.

Capture more than latency. Throughput at that latency, error rate, and the resource utilization that produced it. A flow meeting its latency target at ninety percent CPU is one change away from missing it, and the latency alone does not say that.

Establish it early. A baseline taken after a system has been in production for two years cannot tell you what it used to do, which is exactly the question that matters when someone asks why it feels slow.

On STACKIT. Observability stores the measurements, and its service plans determine how far back a comparison can reach: metrics default to 90 days and can be extended to 26 months. For trend analysis over a system’s life the longer setting is the one that matters, and it is a decision to make before the history you want is gone.

Managed services publish which metrics they expose, for example for PostgreSQL Flex , which is worth reading before designing a baseline around a metric that does not exist. What the platform emits about itself and what your application emits under OPS 7 are separate sources and both belong in the baseline.

Tradeoffs. Cost Optimization. Longer metric retention costs more, and the value only appears when someone asks a question about the past. Retention is the cheapest part of this question and the easiest to cut.

Verify. For your most critical flow, what was its 95th percentile latency six months ago? If you cannot answer, what would you compare today’s measurement against?


PERF 2.2 Measure percentiles rather than averages

Section titled “PERF 2.2 Measure percentiles rather than averages”

Risk if not established: Medium

An average is dominated by the common case and says nothing about the tail. A system averaging 200 milliseconds might have a well-behaved distribution or might serve one request in twenty in four seconds, and those are entirely different systems to the people using them.

Track at least the median, a high percentile such as the 95th, and an extreme such as the 99th. The median tells you about the typical experience, the high percentile about the experience people complain about, and the gap between them tells you whether the system is consistent.

Averages also hide the shape of a regression. A change that makes ten percent of requests much slower barely moves the average and moves the 95th percentile sharply, which is the difference between noticing and not.

Watch the tail specifically for the things that cause it: garbage collection, cache misses, lock contention, retries under REL 5.2, and cold starts. These are invisible in the mean and are most of what users experience as unreliability.

On STACKIT. Percentile calculation is a property of your metric instrumentation rather than of the platform. What you emit under OPS 7.1 determines whether percentiles are available at all: a metric recorded as a single average value cannot be decomposed afterwards.

Tradeoffs. Cost Optimization. Percentile metrics carry more data than a single average, and high-cardinality dimensions multiply that, which is the constraint OPS 7.3 describes.

Verify. For your critical flows, which percentiles are recorded? What is the ratio between the median and the 95th, and has that ratio changed?


PERF 2.3 Compare automatically rather than by periodic review

Section titled “PERF 2.3 Compare automatically rather than by periodic review”

Risk if not established: Medium

A baseline reviewed quarterly detects a regression up to three months after it arrived, by which point the change that caused it is buried among hundreds of others.

Automate the comparison at the two points where it is cheapest. Before merge, where a performance test in the pipeline can catch an obvious regression against a known workload, and where the suspect is one change. In production, comparing against the recent trend, which catches what only appears under real load and real data.

The second is where most regressions are actually found, because the conditions that produce them rarely exist in a test. This is the same argument SEC 1.2 makes about pre-deployment and post-deployment measurement.

Set the threshold from consequence rather than from a round percentage. A flow far from its target can absorb a ten percent regression; one already close to its limit cannot absorb three.

Expect noise. Performance measurements vary between runs, and a comparison without a tolerance produces alerts nobody trusts, which is the fatigue problem from REL 10.2 in a different form.

On STACKIT. Alerting in Observability is where the production-side comparison is expressed. The pipeline-side check runs in STACKIT Pipelines , using whichever load tool you choose, and it is the same mechanism as the health gate in OPS 5.2.

For the pipeline result to mean anything, the environment it runs in has to resemble production in the ways that affect performance, which is OPS 6 and specifically the resource shape from PERF 3.

Tradeoffs. Operational Excellence. Performance tests in a pipeline are slow, and slow pipelines push teams toward batching changes, which works against OPS 4.3. Run the full comparison on a schedule and a fast subset per change.

Verify. How would you find out that your critical flow became 30% slower this week? How long would that take, and what would tell you?


PERF 2.4 Keep the baseline current as the system changes

Section titled “PERF 2.4 Keep the baseline current as the system changes”

Risk if not established: Medium

A baseline that is never updated becomes a historical curiosity. The system it described no longer exists, and comparing against it produces differences that reflect intended changes rather than regressions.

Update it deliberately when something legitimately changes the expected behaviour: a new feature on the path, a deliberate capacity change, a platform version. Record why, so that a future reader can tell an intended shift from an accepted degradation.

The discipline that matters is separating the two. A regression that gets absorbed into the baseline because nobody questioned it is a regression that has become permanent, and the baseline now certifies it. Requiring a reason for every baseline change is what prevents that.

Keep the old values. The interesting question is frequently how the system behaved over years rather than against the last revision, and that history is only available if it was not overwritten.

On STACKIT. The metrics history in Observability is what makes the long comparison possible, subject to the retention decided in PERF 2.1. Where the baseline itself is a recorded artefact rather than a query, it belongs in version control under OPS 3.1, so that changes to it are reviewed like any other.

Tradeoffs. Operational Excellence. Another artefact with an owner and a review, and one that is easy to let drift because nothing breaks when it does.

Verify. When was your baseline last updated, and what changed to justify it? Was any of that change a regression that got accepted rather than fixed?


  • PERF 1 Targets, which the baseline is compared against
  • PERF 7 Evidence-based optimization, which starts from the baseline
  • PERF 9 Performance lifecycle, where accumulated drift is addressed
  • OPS 7 Observability, which supplies the measurements
  • OPS 5.2 Health gates, which use the same comparison mechanism