---
id: PERF06
pillar: performance-efficiency
title: PERF 6. How do you reduce the work the system does?
description: Work not done is free, needs no maintenance and never regresses. Look for it before adding capacity, because scaling hides inefficiency rather than removing it.
status: draft
services: [object-storage, kubernetes-engine]
sidebar:
  order: 15
  label: Reducing work
source_url: "https://framework.stackit.cloud/architecture/pillars/performance-efficiency/perf-06-reduce-work/"
source_file: "docs/architecture/pillars/performance-efficiency/perf-06-reduce-work.mdx"
---

When a system is slow there are two responses: make it do less, or give it more. The second is
faster to implement, costs money permanently, and hides the underlying inefficiency so that the
next enlargement helps less than the last.

Work not done is free. It consumes no capacity, needs no maintenance, and cannot regress. That
makes this the first place to look and the one most often skipped.

## Best practices

- [`PERF 6.1`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-61-remove-unnecessary-work-before-adding-capacity-for-it) Remove unnecessary work before adding capacity for it
- [`PERF 6.2`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-62-cache-what-is-expensive-and-stable-and-decide-the-staleness-deliberately) Cache what is expensive and stable, and decide the staleness deliberately
- [`PERF 6.3`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-63-batch-what-is-chatty-and-compress-what-travels) Batch what is chatty and compress what travels
- [`PERF 6.4`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-64-move-computation-to-the-data-rather-than-data-to-the-computation) Move computation to the data rather than data to the computation

---

## PERF 6.1 Remove unnecessary work before adding capacity for it

**Risk if not established:** Medium

Before optimizing how something is done, ask whether it needs doing. That question is asked less
often than it should be, and its answers are larger than any tuning produces.

The recurring finds: columns selected and never read, payloads serialized and discarded, values
recomputed that have not changed, results fetched that a caller ignores, jobs scheduled whose
output goes nowhere, and validation performed three times on one path because three layers each
defend themselves.

Each of these is invisible in a profile as a problem, because the code doing it is correct and
fast. It shows up as a system doing a lot of work, all of it efficient, most of it pointless.

Start at the top of the stack. An algorithmic improvement or a removed round trip beats a
hardware-level optimization by orders of magnitude, and it makes the system simpler rather than
more complex. That is unusual enough to prioritize.

**On STACKIT.** No platform feature identifies work your system does not need. Distributed tracing
through
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> is what
makes it visible, because a trace shows the calls a request actually made rather than the ones you
believe it makes.

The connection to the platform is in what removal buys. Less work means a smaller machine type
under [`PERF 3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/), fewer nodes, and a lower performance class, all of which are continuing costs
avoided rather than a one-off saving.

**Tradeoffs.** Almost none, which is what makes this the first move. Removing work reduces cost,
consumption and complexity at once, and [`SUS 7`](/architecture/pillars/sustainability/sus-07-efficiency-at-the-top/) and [`COST 3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/) both benefit from the same change.

**Verify.** Take the most expensive request on your critical flow. Which of the work it performs
is consumed by something? What proportion could be removed without any user noticing?

---

## PERF 6.2 Cache what is expensive and stable, and decide the staleness deliberately

**Risk if not established:** Medium

Caching is the most effective and most abused technique in this question. It converts an expensive
repeated operation into a cheap one, and it introduces a second copy of the truth.

Two properties make something worth caching: it is expensive to produce, and it changes less often
than it is read. A value that is cheap to compute gains nothing. A value that changes on every
read gains nothing and adds a consistency problem.

The decision that gets skipped is **staleness**. How out of date may this value be, and what
happens when the origin is unavailable. Both are correctness questions rather than performance
ones, and [`REL 6.2`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-62-serve-reduced-or-stale-results-rather-than-errors-where-correctness-allows) covers the failure case: a cache serving stale data during an outage is
graceful degradation when intended and a correctness bug when accidental.

Invalidation is the hard part and the reason caching earns its reputation. Prefer expiry over
explicit invalidation where the staleness tolerance allows it, because a time-based rule is
self-correcting and an invalidation path that misses a case fails silently and permanently.

Cache at one layer, deliberately. Caches at four layers with different expiry produce behaviour
nobody can reason about during an incident.

**On STACKIT.** Managed options include the <LinkChip href="https://docs.stackit.cloud/products/databases/key-value-store/">Key Value
Store</LinkChip> and
<LinkChip href="https://docs.stackit.cloud/products/databases/redis/">Redis</LinkChip> for application-level caching, and
<LinkChip href="https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/cdn/">CDN
distributions</LinkChip>
for content served at the edge.

Each is a component in the availability composition from [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on) rather than a free addition. A
cache that the flow cannot work without has become a dependency, and its failure behaviour needs
the same treatment as any other.

**Tradeoffs.** **Reliability.** A cache is a dependency and a source of stale answers.
**Security.** A cache holds a copy of the data with its classification, which [`SEC 3.3`](/architecture/pillars/security/sec-03-data-classification/#sec-33-find-the-copies-telemetry-backups-caches-test-data-and-exports) counts.
**Cost Optimization.** Memory and storage traded for latency, and the trade changes as traffic and
data volume change, which is [`PERF 9.3`](/architecture/pillars/performance-efficiency/perf-09-performance-lifecycle/#perf-93-remove-optimizations-whose-justification-has-expired).

**Verify.** For each cache in your workload, what is the maximum staleness, who decided it, and
what happens when the origin is unavailable?

---

## PERF 6.3 Batch what is chatty and compress what travels

**Risk if not established:** Medium

Network round trips have a fixed cost that is independent of payload size. A hundred small
requests cost a hundred latencies; one request carrying the same data costs one.

This is the same pattern [`PERF 4.4`](/architecture/pillars/performance-efficiency/perf-04-data-design/#perf-44-bound-every-result-set-and-eliminate-per-row-round-trips) describes for databases, and it applies equally to service
calls, object storage operations, message publishing and log shipping. The fix is the same: ask
for what you need in one request.

Batching has limits worth respecting. It adds latency for the first item while the batch fills, it
increases the blast radius of a failure, and beyond a certain size it stops helping. It suits
throughput-oriented work and suits interactive requests poorly.

Compression trades CPU for bandwidth and latency. It pays when the data compresses well and the
link is the constraint, and it costs when neither holds. On already-compressed content it is pure
overhead, which is a common accidental configuration.

Note where the saving actually lands. Compressing a payload that crosses zones saves inter-zone
traffic that [`REL 4.1`](/architecture/pillars/reliability/rel-04-redundancy/#rel-41-distribute-compute-across-availability-zones-according-to-the-flows-target) pays for, which is a cost saving as much as a latency one.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/tutorials/optimize-object-storage-performance/">Object Storage performance
guidance</LinkChip>
covers the request-shape decisions that matter for that service, where object size, request rate
and parallelism drive throughput rather than a performance tier.

Cross-zone and cross-region traffic is a real cost as well as a latency cost, which makes locality
a consideration in [`PERF 6.4`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-64-move-computation-to-the-data-rather-than-data-to-the-computation) and a line item in [`COST 6`](/architecture/pillars/cost-optimization/cost-06-data-lifecycle/).

**Tradeoffs.** **Reliability.** A larger batch means more work lost when it fails, and the retry
is more expensive. **Performance Efficiency**, against itself: batching improves throughput and
worsens the latency of the first item, which is the wrong trade for an interactive flow.

**Verify.** For your critical flow, how many separate network operations does one user request
produce? Which of those could be combined?

---

## PERF 6.4 Move computation to the data rather than data to the computation

**Risk if not established:** Medium

Transferring a large data set to filter it somewhere else wastes the transfer. Filtering,
aggregating and projecting at the source moves a small answer instead of a large input.

The principle applies at several scales. A query that returns what is needed rather than
everything and filters in the application. An aggregate computed in the database rather than over
a full result set. Content served from an edge location rather than from the origin. A batch job
running where the data lives rather than pulling it across a boundary.

The counterweight is that the source is frequently the component that cannot scale out, per
[`PERF 5.3`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-53-identify-the-components-that-cannot-scale-out). Pushing work to a saturated database because it is closer to the data makes the
constraint worse. The right answer depends on which side has capacity, which is a measurement
rather than a principle.

Locality also has a cost dimension that is easy to miss. Data crossing a zone or region boundary
is charged as well as slow, so a design that keeps computation and data together is usually
cheaper for the same reason it is faster.

**On STACKIT.** For analytical queries over data held elsewhere,
<LinkChip href="https://docs.stackit.cloud/products/data-and-ai/dremio/">Dremio</LinkChip> is the managed SQL engine with a
unified access layer, which is the shape this best practice describes at the data-platform scale.

At the edge, <LinkChip href="https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/cdn/">CDN
distributions</LinkChip>
serve content closer to the user rather than from the origin.

Placement across zones and regions is bounded by [`SOV 2`](/architecture/pillars/sovereignty/sov-02-placement-and-residency/) and the sovereignty tier from [`SOV 1`](/architecture/pillars/sovereignty/sov-01-sovereignty-tier/). On
STACKIT that constrains less than on platforms with a global footprint, since both regions sit
inside EU jurisdiction, and it is still a constraint to check rather than an optimization to apply
freely.

**Tradeoffs.** **Reliability.** Pushing work to a shared component concentrates load on it, which
is [`PERF 5.3`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-53-identify-the-components-that-cannot-scale-out). **Sovereignty & Compliance.** Moving computation to the data is fine; moving data
to the computation may cross a boundary the classification does not permit.

**Verify.** For your largest data operation, how much data crosses a network boundary and how much
of it is used? Could the filtering happen at the source?

---

## Related

- [`PERF 4.4`](/architecture/pillars/performance-efficiency/perf-04-data-design/#perf-44-bound-every-result-set-and-eliminate-per-row-round-trips) Bounded results and round trips, the same argument applied to data access
- [`PERF 5.3`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-53-identify-the-components-that-cannot-scale-out) Non-scaling components, which limits how much work can be pushed to the source
- [`PERF 7`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/) Evidence-based optimization, which decides where to apply this
- [`REL 6.2`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-62-serve-reduced-or-stale-results-rather-than-errors-where-correctness-allows) Stale results, which caching also produces during failures
- [`SUS 7`](/architecture/pillars/sustainability/sus-07-efficiency-at-the-top/) Efficiency at the top of the stack, the same lever with a different motive
