---
id: PERF02
pillar: performance-efficiency
title: PERF 2. How do you establish a performance baseline and detect regressions?
description: Performance decays through accumulated small increments, none of which anyone would call a regression. Only a baseline makes the accumulation visible.
status: draft
services: [observability]
sidebar:
  order: 11
  label: Baseline
source_url: "https://framework.stackit.cloud/architecture/pillars/performance-efficiency/perf-02-baseline/"
source_file: "docs/architecture/pillars/performance-efficiency/perf-02-baseline.mdx"
---

Systems get slower without anyone breaking anything. Data grows, features accumulate on a common
path, dependencies update, traffic patterns shift. Each increment is negligible and the sum is
not.

A baseline is what converts that from a vague sense that things feel slower into a measurement.
Without one, the first reliable signal is a complaint, by which point the cause is months of
changes rather than one.

## Best practices

- [`PERF 2.1`](/architecture/pillars/performance-efficiency/perf-02-baseline/#perf-21-record-how-the-system-behaves-under-known-conditions) Record how the system behaves under known conditions
- [`PERF 2.2`](/architecture/pillars/performance-efficiency/perf-02-baseline/#perf-22-measure-percentiles-rather-than-averages) Measure percentiles rather than averages
- [`PERF 2.3`](/architecture/pillars/performance-efficiency/perf-02-baseline/#perf-23-compare-automatically-rather-than-by-periodic-review) Compare automatically rather than by periodic review
- [`PERF 2.4`](/architecture/pillars/performance-efficiency/perf-02-baseline/#perf-24-keep-the-baseline-current-as-the-system-changes) Keep the baseline current as the system changes

---

## PERF 2.1 Record how the system behaves under known conditions

**Risk if not established:** Medium

A measurement without its conditions cannot be compared with anything. The baseline needs the
conditions recorded alongside the numbers: what load, what data volume, which version, which
environment, and what else was running.

Baseline the flows from [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume) rather than every endpoint. A baseline covering everything is
expensive to maintain and nobody reads it; one covering the ranked flows gets looked at.

Capture more than latency. Throughput at that latency, error rate, and the resource utilization
that produced it. A flow meeting its latency target at ninety percent CPU is one change away from
missing it, and the latency alone does not say that.

Establish it early. A baseline taken after a system has been in production for two years cannot
tell you what it used to do, which is exactly the question that matters when someone asks why it
feels slow.

**On STACKIT.**
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> stores
the measurements, and its <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/">service
plans</LinkChip>
determine how far back a comparison can reach: metrics default to 90 days and can be extended to
26 months. For trend analysis over a system's life the longer setting is the one that matters, and
it is a decision to make before the history you want is gone.

Managed services publish which metrics they expose, for example <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/observability-metrics-in-postgresql-flex/">for PostgreSQL
Flex</LinkChip>,
which is worth reading before designing a baseline around a metric that does not exist. What the
platform emits about itself and what your application emits under [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) are separate sources and
both belong in the baseline.

**Tradeoffs.** **Cost Optimization.** Longer metric retention costs more, and the value only
appears when someone asks a question about the past. Retention is the cheapest part of this
question and the easiest to cut.

**Verify.** For your most critical flow, what was its 95th percentile latency six months ago? If
you cannot answer, what would you compare today's measurement against?

---

## PERF 2.2 Measure percentiles rather than averages

**Risk if not established:** Medium

An average is dominated by the common case and says nothing about the tail. A system averaging 200
milliseconds might have a well-behaved distribution or might serve one request in twenty in four
seconds, and those are entirely different systems to the people using them.

Track at least the median, a high percentile such as the 95th, and an extreme such as the 99th.
The median tells you about the typical experience, the high percentile about the experience people
complain about, and the gap between them tells you whether the system is consistent.

Averages also hide the shape of a regression. A change that makes ten percent of requests much
slower barely moves the average and moves the 95th percentile sharply, which is the difference
between noticing and not.

Watch the tail specifically for the things that cause it: garbage collection, cache misses, lock
contention, retries under [`REL 5.2`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-52-retry-with-backoff-and-jitter-and-only-what-is-safe-to-retry), and cold starts. These are invisible in the mean and are most
of what users experience as unreliability.

**On STACKIT.** Percentile calculation is a property of your metric instrumentation rather than of
the platform. What you emit under [`OPS 7.1`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-71-emit-metrics-logs-and-traces-and-make-them-joinable) determines whether percentiles are available at all: a
metric recorded as a single average value cannot be decomposed afterwards.

**Tradeoffs.** **Cost Optimization.** Percentile metrics carry more data than a single average,
and high-cardinality dimensions multiply that, which is the constraint [`OPS 7.3`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-73-instrument-for-the-questions-you-cannot-predict) describes.

**Verify.** For your critical flows, which percentiles are recorded? What is the ratio between the
median and the 95th, and has that ratio changed?

---

## PERF 2.3 Compare automatically rather than by periodic review

**Risk if not established:** Medium

A baseline reviewed quarterly detects a regression up to three months after it arrived, by which
point the change that caused it is buried among hundreds of others.

Automate the comparison at the two points where it is cheapest. **Before merge**, where a
performance test in the pipeline can catch an obvious regression against a known workload, and
where the suspect is one change. **In production**, comparing against the recent trend, which
catches what only appears under real load and real data.

The second is where most regressions are actually found, because the conditions that produce them
rarely exist in a test. This is the same argument [`SEC 1.2`](/architecture/pillars/security/sec-01-security-baseline/#sec-12-measure-continuously-and-automatically-rather-than-by-periodic-audit) makes about pre-deployment and
post-deployment measurement.

Set the threshold from consequence rather than from a round percentage. A flow far from its target
can absorb a ten percent regression; one already close to its limit cannot absorb three.

Expect noise. Performance measurements vary between runs, and a comparison without a tolerance
produces alerts nobody trusts, which is the fatigue problem from [`REL 10.2`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-102-alert-on-user-visible-symptoms-rather-than-on-component-metrics-alone) in a different form.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/getting-started/alerting-overview/">Alerting in
Observability</LinkChip>
is where the production-side comparison is expressed. The pipeline-side check runs in <LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/basics/stackit-pipelines/">STACKIT
Pipelines</LinkChip>,
using whichever load tool you choose, and it is the same mechanism as the health gate in `OPS
5.2`.

For the pipeline result to mean anything, the environment it runs in has to resemble production in
the ways that affect performance, which is [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) and specifically the resource shape from
[`PERF 3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/).

**Tradeoffs.** **Operational Excellence.** Performance tests in a pipeline are slow, and slow
pipelines push teams toward batching changes, which works against [`OPS 4.3`](/architecture/pillars/operational-excellence/ops-04-deployment-automation/#ops-43-keep-changes-small-and-deploy-frequently). Run the full
comparison on a schedule and a fast subset per change.

**Verify.** How would you find out that your critical flow became 30% slower this week? How long
would that take, and what would tell you?

---

## PERF 2.4 Keep the baseline current as the system changes

**Risk if not established:** Medium

A baseline that is never updated becomes a historical curiosity. The system it described no longer
exists, and comparing against it produces differences that reflect intended changes rather than
regressions.

Update it deliberately when something legitimately changes the expected behaviour: a new feature
on the path, a deliberate capacity change, a platform version. Record why, so that a future reader
can tell an intended shift from an accepted degradation.

The discipline that matters is separating the two. A regression that gets absorbed into the
baseline because nobody questioned it is a regression that has become permanent, and the baseline
now certifies it. Requiring a reason for every baseline change is what prevents that.

Keep the old values. The interesting question is frequently how the system behaved over years
rather than against the last revision, and that history is only available if it was not
overwritten.

**On STACKIT.** The metrics history in
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/">Observability</LinkChip>
is what makes the long comparison possible, subject to the retention decided in [`PERF 2.1`](/architecture/pillars/performance-efficiency/perf-02-baseline/#perf-21-record-how-the-system-behaves-under-known-conditions). Where
the baseline itself is a recorded artefact rather than a query, it belongs in version control
under [`OPS 3.1`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-31-define-every-production-resource-in-version-control), so that changes to it are reviewed like any other.

**Tradeoffs.** **Operational Excellence.** Another artefact with an owner and a review, and one
that is easy to let drift because nothing breaks when it does.

**Verify.** When was your baseline last updated, and what changed to justify it? Was any of that
change a regression that got accepted rather than fixed?

---

## Related

- [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) Targets, which the baseline is compared against
- [`PERF 7`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/) Evidence-based optimization, which starts from the baseline
- [`PERF 9`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/) Performance lifecycle, where accumulated drift is addressed
- [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/) Observability, which supplies the measurements
- [`OPS 5.2`](/architecture/pillars/operational-excellence/ops-05-safe-deployment/#ops-52-gate-progression-on-health-signals-rather-than-on-elapsed-time) Health gates, which use the same comparison mechanism
