Zum Inhalt springen
Beta

PERF 7. How do you decide what to optimize?

Zuletzt aktualisiert am

Intuition about performance is unreliable in a well-documented way. Systems have too many interacting parts, caches behave unexpectedly, and the code that looks expensive is called once while the trivial function is called a million times.

This is not a failure of skill. It is what happens when a system exceeds what anyone can hold in mind, which is every system worth optimizing.

  • PERF 7.1 Profile to find the constraint before changing anything
  • PERF 7.2 Change one thing and measure whether the constraint moved
  • PERF 7.3 Revert what did not help
  • PERF 7.4 Profile the whole path, not the component you own

PERF 7.1 Profile to find the constraint before changing anything

Section titled “PERF 7.1 Profile to find the constraint before changing anything”

Risk if not established: Medium

An optimization applied without measurement is a guess with a permanent cost: engineering time spent, complexity added, and a benefit nobody verified. Frequently there is no benefit at all, and the complexity stays regardless.

Measure where the time goes rather than where you expect it to. The output you want is an ordered list of where a request spends its time, which makes the constraint obvious rather than debatable.

Profile under realistic conditions. A profile taken at development data volumes measures a different system, as PERF 8.1 explains, and it will point at a different constraint.

Distinguish the two questions a profile can answer. Where does the time go finds the slow part. Which resource is saturated finds the ceiling. They are frequently different, and a component that is slow because it waits on a saturated dependency is not the thing to optimize.

On STACKIT. Distributed tracing through Observability shows where a request spends its time across services, which is the only practical way to answer that question in a distributed system. It depends on the instrumentation and correlation from OPS 7.2: a trace that breaks at a service boundary cannot show you what happens beyond it.

For the resource question, managed services publish which metrics they expose, for example for PostgreSQL Flex . Reading that list before designing a diagnosis is worth the few minutes: a saturation you cannot observe is one you will infer from symptoms.

Language-level profiling inside your process is your own tooling, and it is the layer where PERF 6.1 finds work that should not exist.

Tradeoffs. Performance Efficiency, briefly against itself: profiling in production adds overhead, addressed by sampling rather than by not doing it. Operational Excellence. Instrumentation is design work rather than configuration.

Verify. For your slowest critical flow, where does the time actually go? What measurement produced that answer, and when?


PERF 7.2 Change one thing and measure whether the constraint moved

Section titled “PERF 7.2 Change one thing and measure whether the constraint moved”

Risk if not established: Medium

Changing several things at once produces an improvement nobody can attribute. If two of three changes helped and one hurt, the aggregate looks like a modest win and the harmful change is now permanent.

The loop is simple and rarely followed: measure, form a hypothesis about the constraint, make one change, measure again, compare. The comparison is what turns an opinion into a result.

Expect the constraint to move rather than disappear. Relieving a bottleneck reveals the next one, which is what progress looks like rather than a failure of the exercise. The mistake is assuming the first constraint was the only one and stopping the measurement.

State the expected improvement before making the change. A prediction that turns out wrong is information about the system; a change evaluated only after the fact tends to be judged successful because effort was spent on it.

Watch for improvements that move cost rather than removing it. Caching a slow query makes the flow faster and leaves the slow query, which will surface again the first time the cache misses under load.

On STACKIT. The measurement infrastructure is the same as PERF 2. What matters here is that the before and after are comparable: same data volume, same load, same version of everything else. An environment that differs from production in the ways OPS 6.1 describes will produce a comparison that does not transfer.

Tradeoffs. Operational Excellence. One change at a time is slower than a batch of improvements, and it is the only version that produces knowledge rather than a feeling.

Verify. For your last performance improvement, what was measured before, what was changed, and what was measured after? Was the improvement the one you predicted?


Risk if not established: Medium

Complexity added for a benefit that did not materialize is pure cost, and it is much easier to remove now than in a year when nobody remembers why it exists.

The resistance is not technical. An optimization represents effort, and removing it feels like admitting the effort was wasted. It was, and keeping it wastes more: every future reader has to understand it, every future change has to preserve it, and it will be cited as precedent.

Make the revert a normal outcome rather than an exception. If the measurement in PERF 7.2 does not show the predicted improvement, the change goes back. That expectation is what makes people willing to try things.

Record the attempt even when reverting. A note that caching this value was tried and did not help is worth keeping, because the same idea will occur to someone else, and the second attempt is cheaper to skip than to repeat.

Apply the same discipline to optimizations that worked and have since stopped mattering, which is PERF 9.3.

On STACKIT. No platform feature applies. The revert path is the deployment path from OPS 4.2, which is the same argument for rollback being routine.

Tradeoffs. None. Reverting an unproven change costs the time to revert it and saves everything downstream of keeping it.

Verify. Name an optimization in your codebase whose benefit has been measured. Name one whose benefit has not. Why is the second one still there?


PERF 7.4 Profile the whole path, not the component you own

Section titled “PERF 7.4 Profile the whole path, not the component you own”

Risk if not established: Medium

Teams optimize the component they are responsible for, which is rational and produces local improvements that the user does not experience.

The user experiences the whole path: the browser, the network, the edge, the load balancer, every service hop, the database, and back. A component that accounts for fifteen percent of the total can be made twice as fast and the user notices seven percent.

Start from the end-to-end measurement and decompose it. The largest segment is where to look, and it is frequently outside the boundary of whoever is doing the optimizing, which is the reason this is uncomfortable rather than difficult.

Include the parts that are easy to forget: connection establishment and TLS handshakes, DNS resolution, queueing before the request is picked up, and time spent waiting on a dependency’s retry under REL 5.2.

Where the largest segment belongs to another team or to a third party, that is a finding to escalate rather than to work around. Optimizing your own segment to compensate produces effort with no result.

On STACKIT. End-to-end tracing across services is what makes the decomposition possible, and it requires the correlation identifier from OPS 7.2 to be propagated through every hop, including the asynchronous ones. Where the chain breaks, the segment beyond the break is invisible and will be assumed to be small.

For the client-side segment, which is frequently the largest for a user-facing flow, the platform sees nothing. That measurement comes from your own instrumentation in the client.

Tradeoffs. Operational Excellence. End-to-end tracing across team boundaries requires agreement on the correlation standard, which is OPS 2.1 and an organizational conversation rather than a technical one.

Verify. For your critical flow, what is the end-to-end latency and how does it decompose by segment? Which segment is largest, and who owns it?


  • PERF 2 Baseline, which supplies the before-and-after measurements
  • PERF 6 Reducing work, which is usually what a profile points at
  • PERF 8 Load testing, which produces the realistic conditions to profile under
  • PERF 9.3 Lifecycle, which reverts optimizations that stopped paying
  • OPS 7.2 Correlation, without which the whole-path view does not exist