Skip to content
Beta

PERF 9. How do you maintain performance over the workload's life?

Last updated on

Nothing in this pillar stays done. Data volumes grow and change query plans, features accumulate on common paths, dependencies update, traffic patterns shift as usage matures, and the sizing decision from two years ago is now approximately wrong.

Every increment is individually negligible, which is precisely why nobody is responsible for the sum. This question is the cadence that makes someone responsible.

  • PERF 9.1 Re-examine sizing against measured demand on a cadence
  • PERF 9.2 Re-examine data access as the data grows
  • PERF 9.3 Remove optimizations whose justification has expired
  • PERF 9.4 Treat platform and dependency changes as performance events

PERF 9.1 Re-examine sizing against measured demand on a cadence

Section titled “PERF 9.1 Re-examine sizing against measured demand on a cadence”

Risk if not established: Medium

Sizing decisions decay in both directions. A component sized for a peak that never materialized is over-provisioned; one sized before the workload grew is heading for its ceiling. Neither announces itself.

Set a cadence proportional to how fast demand changes, and give it an owner. Quarterly suits most workloads; a fast-growing one needs more. What matters is that it happens rather than how often.

Compare current utilization against the sizing, and current demand against the ceiling from PERF 3.4. The second is the one that matters more, because it has a lead time: a component at sixty percent of its maximum with demand doubling annually needs a decision now rather than when it arrives.

Pay particular attention to the decisions PERF 3.2 classified as migrations. Those need the most lead time and are the ones a review is most valuable for, because acting late means acting under pressure.

Do it alongside the cost review under COST 9 rather than separately. The two look at the same measurements and reach conclusions that need reconciling, and reconciling them in one conversation is cheaper than in two.

On STACKIT. The measurement history in Observability is what makes a trend visible rather than a snapshot, and the retention decision from PERF 2.1 bounds how long a trend you can see. Extending metric retention is the cheapest preparation for this review and has to happen before the history you want exists.

Tradeoffs. Operational Excellence. A recurring review consumes time from people who could be building. The alternative is discovering the ceiling during an incident.

Verify. When was your sizing last reviewed against measured demand, and what changed as a result? How far is your most constrained component from its ceiling?


PERF 9.2 Re-examine data access as the data grows

Section titled “PERF 9.2 Re-examine data access as the data grows”

Risk if not established: High

Data access degrades non-linearly, which is why it dominates this question. A query that was fine at one volume is not merely slower at ten times the volume; it may be executing a different plan entirely.

Three things to re-examine. Query plans, because an engine’s choice depends on statistics that change as data grows, and a plan that was optimal is not permanent. Indexes, both for the ones now missing on queries that grew important and for the ones now unused and costing writes. Partitioning, because a scheme that distributed evenly may have developed a hot partition as the data skewed.

The trigger is usually growth rather than change. Nobody deployed anything; the table simply crossed a threshold. That makes it invisible to change-based review and detectable only by looking.

Watch the queries whose cost scales with data rather than with usage. A report that scans a full table has a runtime proportional to history, and it will eventually exceed any window it runs in.

On STACKIT. Managed databases publish which metrics they expose, for example for PostgreSQL Flex , and those are what turn this review into a measurement rather than an inspection.

Where the conclusion is that the I/O tier is now too small, PERF 3.2 applies: on managed databases that is a clone to a larger performance class rather than a setting change, which is why noticing early matters more here than on an adjustable axis.

Tradeoffs. Operational Excellence. Reviewing query plans requires someone who can read them, which is a skill that is not evenly distributed and worth deliberately maintaining.

Verify. For your most expensive query, when was its plan last examined? How has its runtime changed over the last year relative to the data volume?


PERF 9.3 Remove optimizations whose justification has expired

Section titled “PERF 9.3 Remove optimizations whose justification has expired”

Risk if not established: Medium

Optimizations accumulate and are almost never removed. Each was justified when added, and the justification is not re-checked, so a system carries complexity for conditions that no longer exist.

The recurring cases: a cache protecting a query that a schema change made fast, a denormalized field whose read path was refactored away, a batch job pre-computing something nobody reads, a connection pool sized for a load pattern that changed, and a read replica added for a report that was retired.

Each costs something permanently: complexity to understand, state to keep consistent, and capacity to run. Removing one is a net gain that nobody is measured on, which is why it does not happen without a deliberate pass.

The prerequisite is having recorded why each optimization exists, which PERF 7.2 asks for. Without that record, removal is a gamble and the safe choice is to keep everything, which is how the accumulation becomes permanent.

Removal is a change like any other and gets the same treatment: measure before, remove, measure after. If performance degrades, the justification still holds and you have learned something worth recording.

On STACKIT. No platform feature applies. What the platform makes visible is the cost side: a cache instance, a replica or a pre-computed data set is a line item that a removal deletes, which is COST 3 reaching the same conclusion from the other direction.

Tradeoffs. Reliability. Removing an optimization risks a regression, which is why it is measured rather than assumed. Operational Excellence. The review costs time and reduces the maintenance burden, which is one of the few places where effort now reduces effort later unambiguously.

Verify. Name three optimizations in your workload. For each, is the condition that justified it still true, and when was that last checked?


PERF 9.4 Treat platform and dependency changes as performance events

Section titled “PERF 9.4 Treat platform and dependency changes as performance events”

Risk if not established: Medium

A version upgrade can change performance in either direction, and nobody is watching for it because the change was operational rather than a feature.

Database major versions change query planners. Runtime upgrades change garbage collection behaviour. Library updates change algorithmic complexity in ways release notes rarely mention. Kubernetes version updates move workloads and reset warm caches.

Treat each as a change that gets the comparison from PERF 2.3 rather than as maintenance that happens quietly. The staged rollout from OPS 5.4 is where a regression should be caught, and it only catches it if somebody is comparing.

Upgrades are also an opportunity worth taking. A new engine version may make a workaround unnecessary, which is a candidate for PERF 9.3, and a query plan that improves is worth knowing about rather than absorbing silently.

Plan for the ones you do not control. Managed service upgrades happen on the provider’s schedule as well as yours, which means the baseline should be watched around those windows rather than only around your own deployments.

On STACKIT. Each managed service publishes a lifecycle reference with dated end-of-support per version, for example for PostgreSQL Flex , which makes a major version change a scheduled event rather than a surprise. OPS 6.3 uses the same dates for environment alignment, and the same upgrade should pass through the same stages.

For clusters, SKE version updates describe the version lifecycle. A cluster upgrade replaces nodes, which resets in-memory state and produces a temporary performance dip that is expected rather than a regression, and distinguishing the two requires knowing the upgrade happened.

Tradeoffs. Operational Excellence. Watching the baseline around every upgrade is more work than upgrading and moving on. It is how a slow accumulation of small regressions gets attributed to something.

Verify. For the last platform or major dependency upgrade, was performance compared before and after? What did that comparison show?


  • PERF 2 Baseline, which supplies the trend this question acts on
  • PERF 3 Selection and sizing, which is re-examined here
  • PERF 7 Evidence-based optimization, which recorded why each optimization exists
  • OPS 6.3 Version alignment, which shares the lifecycle dates
  • COST 9 Review cadence, which should happen in the same conversation