---
id: PERF05
pillar: performance-efficiency
title: PERF 5. How do you scale, and on which signals?
description: Horizontal scaling requires designing for it, and some components never will. The signal you scale on decides whether capacity arrives before or after the peak.
status: draft
services: [kubernetes-engine, compute-engine]
sidebar:
  order: 14
  label: Scaling
source_url: "https://framework.stackit.cloud/architecture/pillars/performance-efficiency/perf-05-scaling/"
source_file: "docs/architecture/pillars/performance-efficiency/perf-05-scaling.mdx"
---

[`REL 7`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/) covers scaling as survival: absorbing demand so the system does not collapse. This
question covers it as throughput: serving more work without each unit of work getting slower.

The mechanisms overlap almost entirely, which is why the two questions cross-reference rather than
repeat. What differs is the target. Reliability asks whether the system stays up; performance asks
whether it stays within [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/).

## Best practices

- [`PERF 5.1`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-51-prefer-horizontal-scaling-which-means-designing-components-to-allow-it) Prefer horizontal scaling, which means designing components to allow it
- [`PERF 5.2`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-52-scale-on-signals-that-predict-saturation-rather-than-confirm-it) Scale on signals that predict saturation rather than confirm it
- [`PERF 5.3`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-53-identify-the-components-that-cannot-scale-out) Identify the components that cannot scale out
- [`PERF 5.4`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-54-account-for-what-scaling-costs-in-time-and-in-money) Account for what scaling costs in time and in money

---

## PERF 5.1 Prefer horizontal scaling, which means designing components to allow it

**Risk if not established:** Medium

Vertical scaling is simpler and has a hard ceiling. Horizontal scaling has a far higher ceiling
and demands properties from the component that have to be designed in.

The properties that decide it: no local state that matters, no fixed identity that peers depend
on, no assumption that one instance sees every request, and tolerance for instances appearing and
disappearing. A component with any of those cannot be scaled out without changing it.

Design them in early where the flow ranking justifies it. Retrofitting statelessness onto a
component that has accumulated local caches, in-memory sessions and background schedulers is a
rewrite rather than a change.

Not everything should scale horizontally. A component with modest demand and no growth is cheaper
and simpler as one instance, and making it distributed adds coordination for no benefit. Apply
this where the ceiling matters.

**On STACKIT.** In
<LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/">Kubernetes Engine</LinkChip>, horizontal
scaling of workloads is the standard Kubernetes mechanism, and node pool capacity is the
cluster-level counterpart, with the limits [`PERF 3.4`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-34-know-where-the-ceiling-of-the-option-you-chose-is) describes.

<LinkChip href="https://docs.stackit.cloud/products/runtime/cloud-foundry/how-tos/use-the-app-autoscaler/">Cloud
Foundry</LinkChip>
provides an application autoscaler, which is the more managed of the two options.

For Compute Engine, scaling out means provisioning instances through the API, CLI or Terraform
provider. There is no managed autoscaling group, so the orchestration is yours to build. That is
the usual division for infrastructure services and it shapes the choice: a workload expected to
scale frequently is cheaper to operate on a runtime that does it for you.

The constraint from [`REL 4.2`](/architecture/pillars/reliability/rel-04-redundancy/#rel-42-make-state-redundant-and-know-where-each-data-set-is-anchored) applies here too. A stateful workload whose volume is bound to one
zone does not become distributable by adding nodes elsewhere.

**Tradeoffs.** **Operational Excellence.** Distributed components introduce coordination, partial
failure and harder debugging. **Cost Optimization.** Many small instances usually cost more than
one large one for the same total capacity, before counting the operational difference.

**Verify.** For each component on your critical flow, can it run as more than one instance? For
those that cannot, what prevents it?

---

## PERF 5.2 Scale on signals that predict saturation rather than confirm it

**Risk if not established:** High

CPU utilization is a lagging indicator. By the time it is high enough to trigger scaling, latency
has already risen, which means the target in [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) was missed before the response began.

Leading indicators predict the saturation instead of reporting it. Queue depth, connection pool
utilization, request rate against known capacity and latency at a high percentile all move before
the system is in trouble. Which one leads depends on where your bottleneck is, which is `PERF
7.1`.

Two configuration properties decide whether it helps. **Bounds**, so that a minimum preserves
headroom and a maximum prevents a runaway from consuming the quota or the budget. And **damping**,
so that the autoscaler does not oscillate. A control loop adding and removing capacity every few
minutes costs more than a fixed size and is less stable.

Scale-in deserves as much attention as scale-out. Aggressive removal during a lull leaves nothing
in place when demand returns, and terminates instances that may be holding work.

**On STACKIT.** The signals come from
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> and
from your instrumentation under [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/). A metric that nobody emits cannot trigger anything, which
is the practical reason [`PERF 2`](/architecture/pillars/performance-efficiency/perf-02-baseline/) comes before this question.

Node pools in <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/getting-started/node-pools/">Kubernetes
Engine</LinkChip>
have configurable minimum and maximum sizes, which are the bounds this best practice asks for at
the cluster level. Workload-level scaling and the signal it uses are yours to configure.

**Tradeoffs.** **Cost Optimization.** A leading indicator scales earlier, which means paying for
capacity slightly before it is needed. That is the point, and it is a real cost.
**Reliability.** An autoscaler responding to an error spike as though it were load will scale into
an outage, which is why the signal choice matters beyond timing.

**Verify.** Which signal triggers scaling for your critical flow? Does it move before or after
latency does?

---

## PERF 5.3 Identify the components that cannot scale out

**Risk if not established:** High

Every system has components that do not scale horizontally, and they determine the ceiling of the
whole flow regardless of how well everything around them scales.

The usual ones: a relational database primary accepting writes, a component holding coordination
state, a licence bound to a host, and anything with a fixed external identity.

Find them explicitly rather than discovering them under load. The test is direct: if demand
doubled, which component could not be given more instances? That list is your architecture's
ceiling.

For each, the options are bounded. Scale it vertically until the largest option runs out. Reduce
the work reaching it, which is [`PERF 6`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/). Partition it, which is [`PERF 4.1`](/architecture/pillars/performance-efficiency/perf-04-data-design/#perf-41-model-for-the-queries-the-workload-actually-issues) and a significant
change. Or accept the ceiling and plan for it under [`PERF 3.4`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-34-know-where-the-ceiling-of-the-option-you-chose-is).

Note that scaling the layers around a non-scaling component makes things worse. Doubling an
application tier in front of a saturated database produces more contention and no throughput,
which is the mistake the fourth design principle names.

**On STACKIT.** Managed databases scale vertically within their flavor and performance class
range, and the performance class change requires a clone per [`PERF 3.2`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-32-establish-which-sizing-decisions-are-reversible-before-you-make-them).

The three-node replica set in <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/basics/architecture-of-postgresql-flex/">PostgreSQL
Flex</LinkChip>
is redundancy and not capacity. Those replicas do not serve reads. It is the single most
consequential fact on this page for a read-heavy workload, because it means the read path and the
write path are bounded by the same instance and a bigger machine is the only lever the service
itself offers.

> From the STACKIT docs: [Architecture › Node level](https://docs.stackit.cloud/products/databases/postgresql-flex/basics/architecture-of-postgresql-flex/#node-level) (Source updated 19.03.2026, copied 05.10.2026)

On the node level, you control the sizing of every node in the instance and its storage. A node is characterized by its number of **CPUs**, its **memory** and the **storage** connected to it. STACKIT calls scaling on this level **vertical scaling**.

STACKIT calls the combination of **CPU** and **memory** flavor. Available flavor sizes are: _Tiny_, _Small_, _Medium_, _Large_, _X‑Large_. Many flavor sizes are available in both **compute optimized** and **memory optimized** variants. Depending on the instance type, you may only choose from a subset of these flavors.

Independently of flavors, you can configure the storage. STACKIT classifies the storage independently in its **size** and its **performance class**. The **performance class** defines the IOPS and the bandwidth of the storage.

Your instance will have a downtime after changing any parameter except upgrading the storage size. You cannot change the storage class/performance class.

Which puts the work back where this best practice says it belongs. Read-heavy load that has
outgrown one instance is answered above the database: a cache in front of it, so that the
repeated reads never arrive; a read model shaped for the query, per [`PERF 4.2`](/architecture/pillars/performance-efficiency/perf-04-data-design/#perf-42-index-deliberately-and-account-for-what-each-index-costs); or the workload
split so that the expensive reads leave the transactional store altogether. All three are design
changes with lead time, which is the argument for knowing this before the growth rather than
during it.

Beyond the largest instance, the answer stops being a bigger machine and becomes a change to the
data architecture. Knowing where that point is, before reaching it, is [`PERF 3.4`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-34-know-where-the-ceiling-of-the-option-you-chose-is).

**Tradeoffs.** Naming a ceiling is uncomfortable and free. Raising one is expensive, which is why
knowing about it early is worth the analysis.

**Verify.** If demand on your critical flow doubled, which component would become the constraint?
What is the plan for that, and how long would it take?

---

## PERF 5.4 Account for what scaling costs in time and in money

**Risk if not established:** Medium

Capacity that arrives after the peak did not help, and capacity that is never released is a
permanent cost that a spike justified once.

The time figure is the whole chain: detecting the condition, deciding, provisioning, starting,
warming caches and connection pools, passing health checks, and receiving traffic. It is almost
always longer than assumed, and node provisioning in particular is minutes rather than seconds.
[`REL 7.4`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-74-account-for-the-time-scaling-takes) covers the same measurement.

That total determines how much standing headroom is needed. Slow scaling combined with thin
headroom is the combination that misses targets; either alone is survivable.

The money figure is easy to lose track of, because autoscaling makes cost a function of traffic
rather than a decision. A maximum set generously to avoid hitting a limit becomes the budget when
something goes wrong, which is why the bound in [`PERF 5.2`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-52-scale-on-signals-that-predict-saturation-rather-than-confirm-it) is a cost control as well as a safety
one.

Reducing startup time pays twice, because it shortens the scaling response and lowers the standing
headroom it requires. Smaller images and less initialization work are [`PERF 6`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/) applied to a place
where it compounds.

**On STACKIT.** Provisioning and startup times depend on machine type, image and what your
workload does on startup, so the honest answer is to measure them in your own environment rather
than adopt a published figure. Measuring once per component and recording it feeds both this
question and [`REL 7.1`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-71-size-for-measured-peaks-and-keep-headroom-for-the-peak-you-did-not-predict).

**Tradeoffs.** **Cost Optimization.** Compensating for slow scaling means more standing capacity,
which is the direct trade. **Sustainability.** That standing capacity is allocated and idle, which
is [`SUS 2`](/architecture/pillars/sustainability/sus-02-right-sizing/).

**Verify.** From the moment load rises, how long until additional capacity serves traffic? What is
the maximum your autoscaler can reach, and what would that cost for a month?

---

## Related

- [`REL 7`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/) Scaling and headroom, the same mechanisms serving availability
- [`PERF 3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/) Selection and sizing, the vertical alternative and its ceiling
- [`PERF 4`](/architecture/pillars/performance-efficiency/perf-04-data-design/) Data design, which is where a non-scaling component is addressed
- [`PERF 6`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/) Reducing work, which raises the ceiling without adding capacity
- [`COST 3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/) Right-sizing, which pulls against standing headroom
