---
id: PERF07
pillar: performance-efficiency
title: PERF 7. How do you decide what to optimize?
description: The bottleneck is almost never where the design discussion assumed. An unmeasured optimization buys permanent complexity for a benefit nobody verified.
status: draft
services: [observability]
sidebar:
  order: 16
  label: Evidence-based optimization
source_url: "https://framework.stackit.cloud/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/"
source_file: "docs/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization.mdx"
---

Intuition about performance is unreliable in a well-documented way. Systems have too many
interacting parts, caches behave unexpectedly, and the code that looks expensive is called once
while the trivial function is called a million times.

This is not a failure of skill. It is what happens when a system exceeds what anyone can hold in
mind, which is every system worth optimizing.

## Best practices

- [`PERF 7.1`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/#perf-71-profile-to-find-the-constraint-before-changing-anything) Profile to find the constraint before changing anything
- [`PERF 7.2`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/#perf-72-change-one-thing-and-measure-whether-the-constraint-moved) Change one thing and measure whether the constraint moved
- [`PERF 7.3`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/#perf-73-revert-what-did-not-help) Revert what did not help
- [`PERF 7.4`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/#perf-74-profile-the-whole-path-not-the-component-you-own) Profile the whole path, not the component you own

---

## PERF 7.1 Profile to find the constraint before changing anything

**Risk if not established:** Medium

An optimization applied without measurement is a guess with a permanent cost: engineering time
spent, complexity added, and a benefit nobody verified. Frequently there is no benefit at all, and
the complexity stays regardless.

Measure where the time goes rather than where you expect it to. The output you want is an ordered
list of where a request spends its time, which makes the constraint obvious rather than debatable.

Profile under realistic conditions. A profile taken at development data volumes measures a
different system, as [`PERF 8.1`](/architecture/pillars/performance-efficiency/perf-08-load-testing/#perf-81-test-with-realistic-data-volume-distribution-and-concurrency) explains, and it will point at a different constraint.

Distinguish the two questions a profile can answer. **Where does the time go** finds the slow
part. **Which resource is saturated** finds the ceiling. They are frequently different, and a
component that is slow because it waits on a saturated dependency is not the thing to optimize.

**On STACKIT.** Distributed tracing through
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> shows
where a request spends its time across services, which is the only practical way to answer that
question in a distributed system. It depends on the instrumentation and correlation from `OPS
7.2`: a trace that breaks at a service boundary cannot show you what happens beyond it.

For the resource question, managed services publish which metrics they expose, for example <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/observability-metrics-in-postgresql-flex/">for
PostgreSQL
Flex</LinkChip>.
Reading that list before designing a diagnosis is worth the few minutes: a saturation you cannot
observe is one you will infer from symptoms.

Language-level profiling inside your process is your own tooling, and it is the layer where
[`PERF 6.1`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-61-remove-unnecessary-work-before-adding-capacity-for-it) finds work that should not exist.

**Tradeoffs.** **Performance Efficiency**, briefly against itself: profiling in production adds
overhead, addressed by sampling rather than by not doing it. **Operational Excellence.**
Instrumentation is design work rather than configuration.

**Verify.** For your slowest critical flow, where does the time actually go? What measurement
produced that answer, and when?

---

## PERF 7.2 Change one thing and measure whether the constraint moved

**Risk if not established:** Medium

Changing several things at once produces an improvement nobody can attribute. If two of three
changes helped and one hurt, the aggregate looks like a modest win and the harmful change is now
permanent.

The loop is simple and rarely followed: measure, form a hypothesis about the constraint, make one
change, measure again, compare. The comparison is what turns an opinion into a result.

Expect the constraint to move rather than disappear. Relieving a bottleneck reveals the next one,
which is what progress looks like rather than a failure of the exercise. The mistake is assuming
the first constraint was the only one and stopping the measurement.

State the expected improvement before making the change. A prediction that turns out wrong is
information about the system; a change evaluated only after the fact tends to be judged successful
because effort was spent on it.

Watch for improvements that move cost rather than removing it. Caching a slow query makes the flow
faster and leaves the slow query, which will surface again the first time the cache misses under
load.

**On STACKIT.** The measurement infrastructure is the same as [`PERF 2`](/architecture/pillars/performance-efficiency/perf-02-baseline/). What matters here is that
the before and after are comparable: same data volume, same load, same version of everything else.
An environment that differs from production in the ways [`OPS 6.1`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-61-let-environments-differ-in-scale-and-data-not-in-shape) describes will produce a
comparison that does not transfer.

**Tradeoffs.** **Operational Excellence.** One change at a time is slower than a batch of
improvements, and it is the only version that produces knowledge rather than a feeling.

**Verify.** For your last performance improvement, what was measured before, what was changed, and
what was measured after? Was the improvement the one you predicted?

---

## PERF 7.3 Revert what did not help

**Risk if not established:** Medium

Complexity added for a benefit that did not materialize is pure cost, and it is much easier to
remove now than in a year when nobody remembers why it exists.

The resistance is not technical. An optimization represents effort, and removing it feels like
admitting the effort was wasted. It was, and keeping it wastes more: every future reader has to
understand it, every future change has to preserve it, and it will be cited as precedent.

Make the revert a normal outcome rather than an exception. If the measurement in [`PERF 7.2`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/#perf-72-change-one-thing-and-measure-whether-the-constraint-moved) does
not show the predicted improvement, the change goes back. That expectation is what makes people
willing to try things.

Record the attempt even when reverting. A note that caching this value was tried and did not help
is worth keeping, because the same idea will occur to someone else, and the second attempt is
cheaper to skip than to repeat.

Apply the same discipline to optimizations that worked and have since stopped mattering, which is
[`PERF 9.3`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/#perf-93-remove-optimizations-whose-justification-has-expired).

**On STACKIT.** No platform feature applies. The revert path is the deployment path from `OPS
4.2`, which is the same argument for rollback being routine.

**Tradeoffs.** None. Reverting an unproven change costs the time to revert it and saves everything
downstream of keeping it.

**Verify.** Name an optimization in your codebase whose benefit has been measured. Name one whose
benefit has not. Why is the second one still there?

---

## PERF 7.4 Profile the whole path, not the component you own

**Risk if not established:** Medium

Teams optimize the component they are responsible for, which is rational and produces local
improvements that the user does not experience.

The user experiences the whole path: the browser, the network, the edge, the load balancer, every
service hop, the database, and back. A component that accounts for fifteen percent of the total
can be made twice as fast and the user notices seven percent.

Start from the end-to-end measurement and decompose it. The largest segment is where to look, and
it is frequently outside the boundary of whoever is doing the optimizing, which is the reason this
is uncomfortable rather than difficult.

Include the parts that are easy to forget: connection establishment and TLS handshakes, DNS
resolution, queueing before the request is picked up, and time spent waiting on a dependency's
retry under [`REL 5.2`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-52-retry-with-backoff-and-jitter-and-only-what-is-safe-to-retry).

Where the largest segment belongs to another team or to a third party, that is a finding to
escalate rather than to work around. Optimizing your own segment to compensate produces effort
with no result.

**On STACKIT.** End-to-end tracing across services is what makes the decomposition possible, and
it requires the correlation identifier from [`OPS 7.2`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-72-propagate-a-correlation-identifier-through-every-hop) to be propagated through every hop,
including the asynchronous ones. Where the chain breaks, the segment beyond the break is invisible
and will be assumed to be small.

For the client-side segment, which is frequently the largest for a user-facing flow, the platform
sees nothing. That measurement comes from your own instrumentation in the client.

**Tradeoffs.** **Operational Excellence.** End-to-end tracing across team boundaries requires
agreement on the correlation standard, which is [`OPS 2.1`](/architecture/pillars/operational-excellence/ops-02-development-standards/#ops-21-agree-the-standards-that-materially-affect-operability-and-write-them-down) and an organizational conversation
rather than a technical one.

**Verify.** For your critical flow, what is the end-to-end latency and how does it decompose by
segment? Which segment is largest, and who owns it?

---

## Related

- [`PERF 2`](/architecture/pillars/performance-efficiency/perf-02-baseline/) Baseline, which supplies the before-and-after measurements
- [`PERF 6`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/) Reducing work, which is usually what a profile points at
- [`PERF 8`](/architecture/pillars/performance-efficiency/perf-08-load-testing/) Load testing, which produces the realistic conditions to profile under
- [`PERF 9.3`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/#perf-93-remove-optimizations-whose-justification-has-expired) Lifecycle, which reverts optimizations that stopped paying
- [`OPS 7.2`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-72-propagate-a-correlation-identifier-through-every-hop) Correlation, without which the whole-path view does not exist
