PERF 6. How do you reduce the work the system does?
Zuletzt aktualisiert am
When a system is slow there are two responses: make it do less, or give it more. The second is faster to implement, costs money permanently, and hides the underlying inefficiency so that the next enlargement helps less than the last.
Work not done is free. It consumes no capacity, needs no maintenance, and cannot regress. That makes this the first place to look and the one most often skipped.
Best practices
Section titled “Best practices”PERF 6.1Remove unnecessary work before adding capacity for itPERF 6.2Cache what is expensive and stable, and decide the staleness deliberatelyPERF 6.3Batch what is chatty and compress what travelsPERF 6.4Move computation to the data rather than data to the computation
PERF 6.1 Remove unnecessary work before adding capacity for it
Section titled “PERF 6.1 Remove unnecessary work before adding capacity for it”Risk if not established: Medium
Before optimizing how something is done, ask whether it needs doing. That question is asked less often than it should be, and its answers are larger than any tuning produces.
The recurring finds: columns selected and never read, payloads serialized and discarded, values recomputed that have not changed, results fetched that a caller ignores, jobs scheduled whose output goes nowhere, and validation performed three times on one path because three layers each defend themselves.
Each of these is invisible in a profile as a problem, because the code doing it is correct and fast. It shows up as a system doing a lot of work, all of it efficient, most of it pointless.
Start at the top of the stack. An algorithmic improvement or a removed round trip beats a hardware-level optimization by orders of magnitude, and it makes the system simpler rather than more complex. That is unusual enough to prioritize.
On STACKIT. No platform feature identifies work your system does not need. Distributed tracing through Observability is what makes it visible, because a trace shows the calls a request actually made rather than the ones you believe it makes.
The connection to the platform is in what removal buys. Less work means a smaller machine type
under PERF 3, fewer nodes, and a lower performance class, all of which are continuing costs
avoided rather than a one-off saving.
Tradeoffs. Almost none, which is what makes this the first move. Removing work reduces cost,
consumption and complexity at once, and SUS 7 and COST 3 both benefit from the same change.
Verify. Take the most expensive request on your critical flow. Which of the work it performs is consumed by something? What proportion could be removed without any user noticing?
PERF 6.2 Cache what is expensive and stable, and decide the staleness deliberately
Section titled “PERF 6.2 Cache what is expensive and stable, and decide the staleness deliberately”Risk if not established: Medium
Caching is the most effective and most abused technique in this question. It converts an expensive repeated operation into a cheap one, and it introduces a second copy of the truth.
Two properties make something worth caching: it is expensive to produce, and it changes less often than it is read. A value that is cheap to compute gains nothing. A value that changes on every read gains nothing and adds a consistency problem.
The decision that gets skipped is staleness. How out of date may this value be, and what
happens when the origin is unavailable. Both are correctness questions rather than performance
ones, and REL 6.2 covers the failure case: a cache serving stale data during an outage is
graceful degradation when intended and a correctness bug when accidental.
Invalidation is the hard part and the reason caching earns its reputation. Prefer expiry over explicit invalidation where the staleness tolerance allows it, because a time-based rule is self-correcting and an invalidation path that misses a case fails silently and permanently.
Cache at one layer, deliberately. Caches at four layers with different expiry produce behaviour nobody can reason about during an incident.
On STACKIT. Managed options include the Key Value Store and Redis for application-level caching, and CDN distributions for content served at the edge.
Each is a component in the availability composition from REL 1.3 rather than a free addition. A
cache that the flow cannot work without has become a dependency, and its failure behaviour needs
the same treatment as any other.
Tradeoffs. Reliability. A cache is a dependency and a source of stale answers.
Security. A cache holds a copy of the data with its classification, which SEC 3.3 counts.
Cost Optimization. Memory and storage traded for latency, and the trade changes as traffic and
data volume change, which is PERF 9.3.
Verify. For each cache in your workload, what is the maximum staleness, who decided it, and what happens when the origin is unavailable?
PERF 6.3 Batch what is chatty and compress what travels
Section titled “PERF 6.3 Batch what is chatty and compress what travels”Risk if not established: Medium
Network round trips have a fixed cost that is independent of payload size. A hundred small requests cost a hundred latencies; one request carrying the same data costs one.
This is the same pattern PERF 4.4 describes for databases, and it applies equally to service
calls, object storage operations, message publishing and log shipping. The fix is the same: ask
for what you need in one request.
Batching has limits worth respecting. It adds latency for the first item while the batch fills, it increases the blast radius of a failure, and beyond a certain size it stops helping. It suits throughput-oriented work and suits interactive requests poorly.
Compression trades CPU for bandwidth and latency. It pays when the data compresses well and the link is the constraint, and it costs when neither holds. On already-compressed content it is pure overhead, which is a common accidental configuration.
Note where the saving actually lands. Compressing a payload that crosses zones saves inter-zone
traffic that REL 4.1 pays for, which is a cost saving as much as a latency one.
On STACKIT. Object Storage performance guidance covers the request-shape decisions that matter for that service, where object size, request rate and parallelism drive throughput rather than a performance tier.
Cross-zone and cross-region traffic is a real cost as well as a latency cost, which makes locality
a consideration in PERF 6.4 and a line item in COST 6.
Tradeoffs. Reliability. A larger batch means more work lost when it fails, and the retry is more expensive. Performance Efficiency, against itself: batching improves throughput and worsens the latency of the first item, which is the wrong trade for an interactive flow.
Verify. For your critical flow, how many separate network operations does one user request produce? Which of those could be combined?
PERF 6.4 Move computation to the data rather than data to the computation
Section titled “PERF 6.4 Move computation to the data rather than data to the computation”Risk if not established: Medium
Transferring a large data set to filter it somewhere else wastes the transfer. Filtering, aggregating and projecting at the source moves a small answer instead of a large input.
The principle applies at several scales. A query that returns what is needed rather than everything and filters in the application. An aggregate computed in the database rather than over a full result set. Content served from an edge location rather than from the origin. A batch job running where the data lives rather than pulling it across a boundary.
The counterweight is that the source is frequently the component that cannot scale out, per
PERF 5.3. Pushing work to a saturated database because it is closer to the data makes the
constraint worse. The right answer depends on which side has capacity, which is a measurement
rather than a principle.
Locality also has a cost dimension that is easy to miss. Data crossing a zone or region boundary is charged as well as slow, so a design that keeps computation and data together is usually cheaper for the same reason it is faster.
On STACKIT. For analytical queries over data held elsewhere, Dremio is the managed SQL engine with a unified access layer, which is the shape this best practice describes at the data-platform scale.
At the edge, CDN distributions serve content closer to the user rather than from the origin.
Placement across zones and regions is bounded by SOV 2 and the sovereignty tier from SOV 1. On
STACKIT that constrains less than on platforms with a global footprint, since both regions sit
inside EU jurisdiction, and it is still a constraint to check rather than an optimization to apply
freely.
Tradeoffs. Reliability. Pushing work to a shared component concentrates load on it, which
is PERF 5.3. Sovereignty & Compliance. Moving computation to the data is fine; moving data
to the computation may cross a boundary the classification does not permit.
Verify. For your largest data operation, how much data crosses a network boundary and how much of it is used? Could the filtering happen at the source?
Related
Section titled “Related”PERF 4.4Bounded results and round trips, the same argument applied to data accessPERF 5.3Non-scaling components, which limits how much work can be pushed to the sourcePERF 7Evidence-based optimization, which decides where to apply thisREL 6.2Stale results, which caching also produces during failuresSUS 7Efficiency at the top of the stack, the same lever with a different motive