Skip to content
Beta

REL 7. How do you scale to absorb demand, and how much headroom do you keep?

Last updated on

Capacity belongs to this pillar because saturation is a failure mode. A system that cannot absorb its peak does not slow down gracefully; it queues, times out, retries, and collapses, which is REL 6.3 seen from the other side.

The reliability question is narrower than the performance one in PERF 5. Here it is not about being fast. It is about not falling over, and about knowing where the ceiling is before you reach it under pressure.

  • REL 7.1 Size for measured peaks and keep headroom for the peak you did not predict
  • REL 7.2 Automate scaling where the workload permits it
  • REL 7.3 Find every scaling limit before you approach it
  • REL 7.4 Account for the time scaling takes

REL 7.1 Size for measured peaks and keep headroom for the peak you did not predict

Section titled “REL 7.1 Size for measured peaks and keep headroom for the peak you did not predict”

Risk if not established: High

Sizing to the average guarantees the peak fails. Sizing to the observed peak leaves nothing for the peak that has not happened yet, and peaks are not drawn from a stable distribution: a marketing campaign, a public holiday, a retry storm from a client you do not control, or a neighbouring system’s failure can all move the ceiling.

Headroom is the margin between normal operation and the point where behaviour degrades. How much you need depends on how quickly you can add capacity, which is REL 7.4, and on the consequence of being wrong, which is the flow ranking from REL 2.2. A flow that can scale in ninety seconds needs less standing headroom than one that requires a twenty-minute provisioning cycle.

Measure the peak rather than estimating it. Peaks are frequently invisible in averaged metrics: a minute-level average hides a five-second burst that saturated a connection pool, and the incident report will say the system was at forty percent utilization.

Include the failure case in the sizing. When one zone is lost, the remaining zones absorb its traffic. A two-zone deployment running at sixty percent in each zone has no capacity to survive losing one, which makes the redundancy from REL 4.1 decorative.

On STACKIT. Sizing decisions are yours; the platform’s contribution is stating what the ceilings are.

In Kubernetes Engine , SKE reserves system resources on every node, so allocatable capacity is below the machine type’s nominal figure. The reservation is tiered by node size, which means small nodes lose a larger proportion than large ones. Sizing against nominal rather than allocatable capacity is a common way to be quietly ten percent short.

Tradeoffs. Cost Optimization. Headroom is capacity that produces nothing until it is needed, and it is the first thing a cost review finds. Sustainability. The same capacity is allocated and idle, which is the divergence described in SUS 4. Both are answered by the stated target from REL 1: headroom sized to an agreed target is justified, headroom sized to anxiety is not.

Verify. What is your measured peak, at what resolution was it measured, and what is your current headroom above it? If one availability zone were lost right now, would the remaining capacity carry the load?


REL 7.2 Automate scaling where the workload permits it

Section titled “REL 7.2 Automate scaling where the workload permits it”

Risk if not established: Medium

Manual scaling has a response time measured in human availability, which means the peak is over before anyone reacts, or nobody was awake.

Automation only works for workloads that can absorb it. Components that hold local state, take minutes to become ready, or hold long-lived connections resist scaling out, and forcing it produces a system that thrashes. Establish which of your components are actually elastic before designing around the assumption that they all are.

Two configuration points decide whether autoscaling helps or hurts:

The signal. CPU is a lagging indicator: by the time it is saturated, latency has already risen. Queue depth, connection count or request rate usually predict saturation earlier and produce scaling that arrives in time.

The bounds and the damping. A minimum that maintains headroom, a maximum that prevents a runaway from consuming the budget or the quota, and cooldown behaviour that stops oscillation. An autoscaler that adds and removes capacity every few minutes costs more than a fixed size and is less stable.

Scaling down deserves as much thought as scaling up. Aggressive scale-in during a lull leaves nothing in place when demand returns, and it terminates instances that may be holding work.

On STACKIT. Kubernetes Engine node pools have configurable minimum and maximum sizes, which is the cluster-level lever; workload scaling inside the cluster uses the standard Kubernetes mechanisms you configure yourself.

Cloud Foundry provides an App Autoscaler for applications running on it, which is the more managed option of the two.

For Compute Engine, scaling is provisioning: creating and deleting VMs through the API, CLI or Terraform provider. There is no managed autoscaling group, so the orchestration is yours to build. That is the usual division for infrastructure services, and a design that assumes otherwise finds out late.

Tradeoffs. Operational Excellence. An autoscaler is a control loop with thresholds to tune and failure modes of its own, including scaling into an outage because errors looked like load.

Verify. Which components scale automatically, on which signal, and what are the bounds? When did the autoscaler last act, and was the action correct?


REL 7.3 Find every scaling limit before you approach it

Section titled “REL 7.3 Find every scaling limit before you approach it”

Risk if not established: High

Every system has a ceiling. The question is whether you find it during capacity planning or during an incident.

Limits come in three kinds and they are usually discovered in the wrong order:

Platform quotas. Documented, and the easiest to plan around once you know they exist.

Architectural limits. A single database that cannot be sharded, a component that only runs as one instance, a licence tied to a host. These bind far below any quota and take months to remove.

Emergent limits. Connection pools, file descriptors, thread pools, IP address space. Nobody configured these as capacity decisions and they bind at values nobody chose.

Find the binding one, because it is the only number that matters. Relieving it will reveal the next, which is what progress looks like rather than a failure of the exercise.

On STACKIT. SKE’s limits are stated in quotas and limits .

From the STACKIT docsQuotas and limits › Number of nodes in a SKE clusterSource updated 28.09.2026 · copied 05.10.2026

The maximum number of running nodes in a cluster is limited to 1000, regardless of the configured size of your node pools.

This restriction is enforced by the cluster autoscaler, which uses the --max-nodes-total flag to set the upper limit for your cluster.

If you created your SKE Cluster in a SNA network, the maximum node count might be even lower due to limitations in the available address space of that network.

When specifying node pools for your SKE Cluster, the following restrictions apply:

  • The maximum of one individual node pool must always 1000 or less.
  • The minimum of one individual node pool must always 1000 or less.
  • The total of all maximum sizes can be greater than 1000.
  • The total of all minimum sizes must always 1000 or less.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

One caveat belongs in the sizing conversation: clusters in SNA networks may have lower maximum node counts because of address space. That is an emergent limit stated up front rather than one discovered under load.

Project quotas bound consumption per project, though they cover IaaS and Cloud Foundry resources rather than everything in a project. Treat quota errors as capacity signals rather than transient faults: retrying against a quota, as noted in REL 5.2, consumes rate without ever succeeding.

Tradeoffs. Little beyond the analysis time. The main cost is discovering an architectural limit that requires real work to remove, which is unwelcome and much cheaper to learn now.

Verify. What is the binding constraint on your critical flow’s throughput, what is its value, and how far is current peak from it? If the answer is a platform quota, what is the architectural limit behind it?


REL 7.4 Account for the time scaling takes

Section titled “REL 7.4 Account for the time scaling takes”

Risk if not established: Medium

Capacity that arrives after the peak did not help. The relevant figure is not whether you can scale but how long it takes from the signal to serving traffic, and it is almost always longer than assumed.

Count the whole chain: detecting the condition, deciding to act, provisioning the resource, starting it, warming caches and connection pools, passing health checks, and receiving traffic. Node provisioning in particular is minutes rather than seconds, so a cluster autoscaler responding to sudden load will be late.

That total is what determines your standing headroom in REL 7.1. Slow scaling and thin headroom is the combination that fails; either one alone is survivable.

Where scaling is too slow for the demand shape, the alternatives are to pre-scale against a known event, to shed load under REL 6.3, or to accept the degradation. All three are legitimate and all three are better than an autoscaler nobody timed.

On STACKIT. Timings depend on machine type, image, and what your workload does on startup, so the honest answer is to measure them in your own environment rather than take a published figure.

The measurement is worth doing once per component and recording, because it feeds directly into REL 7.1 and into the RTO in REL 1.2.

There is no service level commitment on provisioning time either, so there is no figure to plan against and nobody to hold to one. That is what turns the measurement above from diligence into the only number you have.

Tradeoffs. Cost Optimization. Compensating for slow scaling means more standing capacity, which is the direct trade. Faster startup, through smaller images and less initialization work, reduces both, which is one of the places PERF 6 pays for itself twice.

Verify. Time it: from the moment load rises, how long until additional capacity is serving traffic? Is that shorter than the time your headroom buys you?


  • REL 1 Reliability targets, which decide how much headroom is justified
  • REL 4 Redundancy, whose failure case must be included in the sizing
  • REL 6 Graceful degradation, which is what happens when scaling is not enough
  • PERF 5 Scaling, the same mechanism serving throughput rather than survival
  • COST 3 Right-sizing, and SUS 2, which both pull against standing headroom