Skip to content
Beta

PERF 8. How do you test performance under realistic load?

Last updated on

A load test answers a question that production will otherwise answer for you, at a worse moment and in front of users.

The value depends almost entirely on how closely the test resembles reality. A test with a thousand tidy rows and uniform traffic will pass, and it will have exercised none of the behaviour that matters: the query plan at scale, the connection pool under concurrency, the cache at a realistic hit rate, the collapse past saturation.

  • PERF 8.1 Test with realistic data volume, distribution and concurrency
  • PERF 8.2 Include the region past saturation, because that behaviour is a design property
  • PERF 8.3 Test where the result transfers to production
  • PERF 8.4 Keep testing after launch

PERF 8.1 Test with realistic data volume, distribution and concurrency

Section titled “PERF 8.1 Test with realistic data volume, distribution and concurrency”

Risk if not established: High

Three dimensions decide whether a test measures the same system production runs, and all three are routinely wrong.

Volume changes query plans. An index scan at ten thousand rows becomes a sequential scan at ten million, and the engine chooses differently. A test at development volume measures a different execution path, not a faster version of the same one.

Distribution is where the edge cases live. Real data is skewed: a few customers with a hundred times the average, records with unexpected nulls, legacy rows in a shape the current code barely handles. Uniformly generated data exercises the happy path at scale, which is the least interesting thing to know.

Concurrency is what produces contention. Connection pool exhaustion, lock waits, and cache stampedes only appear when requests overlap, and a sequential test at high volume finds none of them.

Traffic shape matters as much as traffic volume. A steady rate and a burst averaging the same rate produce different behaviour, and the burst is what saturates a pool.

Do not use production data to obtain realism. OPS 6.4 and SEC 3.3 both explain why, and generated data preserving volume, cardinality and skew is the workable middle.

On STACKIT. The environment has to resemble production in the dimensions that affect performance, which is OPS 6.1 and specifically the resource shape from PERF 3.3. A test against a single database instance where production runs a replica set, or against a different performance class, produces a number that does not transfer.

Managed database flavors and performance classes make this concrete: the I/O ceiling differs by an order of magnitude across classes, so a test at a lower class measures a different constraint than production has.

Tradeoffs. Cost Optimization. A realistic test environment costs close to production while it exists, which is the argument for creating it on demand under OPS 3 and destroying it afterwards under SUS 6.

Verify. How much data does your load test run against, and how does its distribution compare with production? At what concurrency does it run?


PERF 8.2 Include the region past saturation, because that behaviour is a design property

Section titled “PERF 8.2 Include the region past saturation, because that behaviour is a design property”

Risk if not established: High

Most load tests stop at the target and report success. What happens beyond it is left untested, which means the system’s behaviour under overload is whatever emerges rather than what was designed.

That behaviour is rarely graceful on its own. Queues grow, latency rises, timeouts fire, retries add load, and throughput collapses below what the system could have sustained. REL 6.3 describes the mechanism and the deliberate alternative.

Test past the target deliberately to find three things. The saturation point, which is the actual capacity rather than the assumed one. The behaviour beyond it, which shows whether load shedding works or whether the system collapses. And the recovery, which is whether it returns to normal when load drops or stays degraded.

The last is the one that surprises people. A system that recovers only after a restart has a failure mode that a test stopping at the target would never reveal.

Note where the saturation point sits relative to your headroom in REL 7.1. Those two numbers together are the capacity plan; either alone is not.

On STACKIT. Testing to saturation means reaching the ceilings from PERF 3.4, including platform quotas. A test that hits a quota has found a real limit rather than a failed test, and distinguishing a quota error from a capacity error is worth doing deliberately, per REL 5.2.

Run it in an isolated project so that saturating a shared component does not affect anything else, which is the segmentation from SEC 2.1 serving a different purpose.

Tradeoffs. Reliability. Deliberately saturating a system is disruptive and, if the environment shares anything with production, risky. That is an argument for isolation rather than for skipping it.

Verify. At what load does your system stop meeting its target, and what does it do at twice that? Does it recover on its own when the load stops?


PERF 8.3 Test where the result transfers to production

Section titled “PERF 8.3 Test where the result transfers to production”

Risk if not established: Medium

A load test result that does not predict production behaviour is worse than no result, because it produces confidence.

The differences that break transferability are the same ones OPS 6.1 names: different topology, different resource shape, different platform versions, a stand-in component instead of the real one. A test against an in-memory queue tells you nothing about the real broker, and one against a single node tells you nothing about coordination.

Scale differences are acceptable and need accounting. A test at a quarter of production capacity can be informative if you know which relationship holds. Some things scale linearly, some do not, and assuming the first without checking is where extrapolation goes wrong.

Record what the test environment did not cover. An explicit list of untested behaviour is useful; an implicit one is discovered in production.

Where a full-fidelity test is genuinely impractical, testing in production with a bounded blast radius is a legitimate alternative, and it is REL 10.3 applied to load rather than to faults.

On STACKIT. Creating a representative environment from the definitions in OPS 3 is what makes fidelity affordable, since the alternative is either a permanent second environment or a test that does not transfer.

The specific things to match are the ones that determine the constraint: the machine type family from PERF 3.3, the database flavor and performance class, and the storage class from PERF 4.3. Matching the size while missing the class is the common error, and it changes which resource saturates first.

Tradeoffs. Cost Optimization. Fidelity costs money in direct proportion. Matching the shape while reducing the scale is usually the best available compromise.

Verify. List the ways your test environment differs from production. For each, does it change which resource saturates first?


Risk if not established: Medium

A load test before launch establishes what the system could do on that day, against that data, with that code. All three drift, and the drift is exactly what PERF 2 is designed to detect and this question is designed to quantify.

Production conditions move away from any single test: data grows, traffic patterns shift as usage matures, dependencies update, and features accumulate on the common path. A capacity figure from launch is a historical fact within a year.

Re-run on a cadence and after material change. The cadence catches accumulation; the trigger catches the specific change that moved the ceiling.

Combine it with the production baseline rather than treating the two as alternatives. The baseline tells you what the system does under real load; the test tells you what it would do under more. Only the second answers a capacity question, and only the first reflects reality.

Watch for the case where the test still passes and production has degraded. That usually means the test data has stopped resembling production data, which is PERF 8.1 decaying quietly.

On STACKIT. Environments created on demand from OPS 3 make a recurring test affordable, and destroying them afterwards is SUS 6.

The measurement side is the same Observability history as PERF 2, where the retention setting decides how far back a capacity trend can be compared.

Tradeoffs. Cost Optimization. A recurring load test consumes environment cost and engineering time on a schedule, producing nothing a customer sees, which is the shape most of Operational Excellence has.

Verify. When was your last load test, against what data volume, and how does that volume compare with production today?


  • PERF 1 Targets, which the test is measured against
  • PERF 2 Baseline, the production-side counterpart to this test
  • PERF 3.4 Ceilings, which a saturation test finds empirically
  • REL 6.3 Load shedding, whose behaviour past saturation this verifies
  • OPS 6 Environment consistency, which decides whether the result transfers