---
id: SUS04
pillar: sustainability
title: SUS 4. How do you increase utilization density?
description: A server at fifteen percent utilization does not consume fifteen percent of the resources. Idle infrastructure draws power and occupies hardware already built.
status: draft
services: [kubernetes-engine]
sidebar:
  order: 13
  label: Utilization density
source_url: "https://framework.stackit.cloud/architecture/pillars/sustainability/sus-04-utilization-density/"
source_file: "docs/architecture/pillars/sustainability/sus-04-utilization-density.mdx"
---

Where the provider layer is already efficient, as [`SUS 1.3`](/architecture/pillars/sustainability/sus-01-measurement/#sus-13-separate-what-the-provider-controls-from-what-you-control) establishes for STACKIT, the variable
left to you is how much infrastructure you occupy and how much useful work happens per unit of
it.

The reason density dominates is that consumption is not proportional to utilization. A lightly
used server draws substantial power, occupies physical space, and represents hardware whose
manufacturing footprint was paid regardless of what it subsequently does. Consolidating the same
work onto fewer, better-used resources reduces consumption far more than making an individual
workload marginally more efficient.

## Best practices

- [`SUS 4.1`](/architecture/pillars/sustainability/sus-04-utilization-density/#sus-41-consolidate-onto-fewer-better-utilized-resources) Consolidate onto fewer, better-utilized resources
- [`SUS 4.2`](/architecture/pillars/sustainability/sus-04-utilization-density/#sus-42-prefer-active-active-over-idle-standby) Prefer active-active over idle standby
- [`SUS 4.3`](/architecture/pillars/sustainability/sus-04-utilization-density/#sus-43-account-for-what-the-platform-reserves-before-it-reaches-your-workload) Account for what the platform reserves before it reaches your workload
- [`SUS 4.4`](/architecture/pillars/sustainability/sus-04-utilization-density/#sus-44-know-where-density-conflicts-with-isolation-and-decide-rather-than-default) Know where density conflicts with isolation, and decide rather than default

---

## SUS 4.1 Consolidate onto fewer, better-utilized resources

**Risk if not established:** Medium

Estates fragment over time. Each workload gets its own instance because that was simplest at the
moment it was created, and the result is many resources each doing a little.

Consolidation is the mechanical answer: the same work on fewer, larger, better-used resources. It
reduces the number of things drawing power, the number of things carrying a system overhead, and
the number of things to operate, which is [`OPS 10`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/) benefiting from the same change.

Two properties make a workload a consolidation candidate. It tolerates neighbours, meaning it does
not need dedicated hardware for isolation or for predictable performance. And its demand profile
complements others, so peaks do not coincide.

The second is the one that decides whether consolidation helps. Two workloads peaking at the same
hour need the sum of their peaks whether they share a host or not; two peaking at different times
need only the larger.

**On STACKIT.** Container platforms are the practical consolidation mechanism. In <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/">Kubernetes
Engine</LinkChip>, many workloads share a
node pool, and how densely depends on the resource requests you set: requests that are generous
relative to actual use reserve capacity that nothing consumes, which reproduces the fragmentation
problem inside the cluster.

That makes container resource requests the same decision as instance sizing in [`SUS 2`](/architecture/pillars/sustainability/sus-02-right-sizing/), applied at
a finer grain and with the same tendency to round up.

**Tradeoffs.** **Reliability.** Consolidation concentrates blast radius: one host or one cluster
failing affects more workloads. **Performance Efficiency.** Neighbours contend for shared
resources, which makes performance less predictable.

**Verify.** How many separate compute resources does your estate run, and what is the average
utilization across them? How many could be combined?

---

## SUS 4.2 Prefer active-active over idle standby

**Risk if not established:** Medium

This is the sharpest conflict between this pillar and Reliability, and it does not resolve in this
pillar's favour. Redundancy is deliberately reserved idle capacity, and [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/) is not waived for
environmental reasons.

What is available is the choice of shape. An active-passive design keeps a standby doing nothing
until a failure that may never occur. An active-active design uses the same total capacity to
serve traffic, so the redundancy is a property of how the load is spread rather than a
reservation.

Where the workload allows it, active-active gives the same failure tolerance with the capacity
doing useful work. Where it does not, because the component cannot run as multiple active
instances under [`PERF 5.3`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-53-identify-the-components-that-cannot-scale-out), standby is the correct answer and the idle capacity is justified.

Where standby is unavoidable, two reductions remain. Keep it minimal and scale it on failover
rather than mirroring production continuously, and check that the standby is actually needed at
the size it is, which is [`SUS 2.3`](/architecture/pillars/sustainability/sus-02-right-sizing/#sus-23-distinguish-headroom-from-slack) distinguishing headroom from slack.

**On STACKIT.** The topology choice appears directly in the availability figures from [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on),
where a system group spanning two zones is committed to a higher availability than a single
instance. What that comparison does not say is whether both instances serve traffic, which is your
design decision and the one this best practice is about.

For managed databases the replica topology is fixed by the service, so the shape is not yours to
choose, and the replicas do not serve reads. That capacity waits rather than works, which is the
fact [`PERF 5.3`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-53-identify-the-components-that-cannot-scale-out) starts from.

**Tradeoffs.** **Reliability.** Active-active is harder to reason about, requires the workload to
tolerate concurrent instances, and has failure modes that active-passive does not. It is the right
default where it fits and not a universal improvement.

**Verify.** For each redundant component, is the redundant capacity serving traffic or waiting?
For those waiting, what prevents them from serving?

---

## SUS 4.3 Account for what the platform reserves before it reaches your workload

**Risk if not established:** Medium

Allocated capacity and usable capacity are not the same number. Every layer between the hardware
and your process takes something: the hypervisor, the operating system, the orchestrator, the
agents.

Sizing against the nominal figure means the workload has less than intended, which produces either
a performance problem or a compensating over-allocation. Both are worse than knowing the
reservation.

The reservation is frequently not proportional. Where a fixed overhead exists per unit, many small
units carry more total overhead than fewer large ones for the same nominal capacity, which is an
argument for larger units that runs alongside the consolidation argument in [`SUS 4.1`](/architecture/pillars/sustainability/sus-04-utilization-density/#sus-41-consolidate-onto-fewer-better-utilized-resources).

**On STACKIT.** Kubernetes Engine documents this precisely, and it is the clearest example of the
non-proportionality. The <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/operations/quotas-and-limits/">quotas and
limits</LinkChip>
page describes system resource reservations on every node, tiered so that a larger share of a
small node is reserved than of a large one.

The consequence is direct: a node pool of many small nodes loses proportionally more capacity to
system overhead than a pool of fewer large nodes providing the same nominal total. That is a
density decision hiding inside a node size decision, and it is not visible from the machine type
figures.

The same page notes that clusters in certain network configurations have lower node maxima because
of address space, which is a different constraint and belongs in [`PERF 3.4`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-34-know-where-the-ceiling-of-the-option-you-chose-is).

**Tradeoffs.** **Reliability.** Fewer, larger nodes means a larger blast radius per node failure
and coarser granularity when scaling, which is the same trade [`SUS 4.1`](/architecture/pillars/sustainability/sus-04-utilization-density/#sus-41-consolidate-onto-fewer-better-utilized-resources) makes.

**Verify.** For your cluster, what is the difference between nominal and allocatable capacity? How
does that proportion change with node size?

---

## SUS 4.4 Know where density conflicts with isolation, and decide rather than default

**Risk if not established:** Medium

Density and isolation pull against each other directly. Every boundary that [`SEC 2.1`](/architecture/pillars/security/sec-02-segmentation/#sec-21-choose-the-isolation-strength-from-the-protection-need-not-from-a-default) describes
prevents consolidation across it, and every consolidation removes a boundary.

That is a real tradeoff rather than a problem to solve. A workload whose classification requires a
separate project or a dedicated instance is not a density candidate, and pursuing density there
would be optimizing the wrong thing.

What this best practice asks for is that the boundary be a decision rather than a default. A
workload running alone because its classification requires it is correct. A workload running alone
because it was created that way is a consolidation candidate that nobody examined.

Use the classification from [`SEC 3`](/architecture/pillars/security/sec-03-data-classification/) and the tier from [`SOV 1`](/architecture/pillars/sovereignty/sov-01-sovereignty-tier/) to sort them. Where the protection
need permits shared infrastructure with logical separation, the mechanisms in [`SEC 2.1`](/architecture/pillars/security/sec-02-segmentation/#sec-21-choose-the-isolation-strength-from-the-protection-need-not-from-a-default) provide it
and the density is available. Where it does not, the capacity is justified.

**On STACKIT.** The in-project separation mechanisms from [`SEC 2.1`](/architecture/pillars/security/sec-02-segmentation/#sec-21-choose-the-isolation-strength-from-the-protection-need-not-from-a-default) are what make density and
isolation partly compatible: namespaces with Kubernetes RBAC and network policy, KMS key rings,
per-service access models. Each allows workloads to share infrastructure while remaining
separated, with the caveat recorded there about who can bypass those boundaries.

Where the requirement is hardware-level isolation, for example under a classification that demands
it, protecting data in use carries a consumption premium as well as a cost one, per [`SOV 5`](/architecture/pillars/sovereignty/sov-05-operator-access/). That
is a justified premium rather than waste, and worth recording as such so a later density review
does not treat it as a target.

**Tradeoffs.** **Security** and **Sovereignty & Compliance**, both directly. This best practice
exists to make the tradeoff explicit rather than to resolve it in either direction.

**Verify.** For each workload running on dedicated infrastructure, what requires that? How many
are isolated by decision rather than by history?

---

## Related

- [`SUS 2`](/architecture/pillars/sustainability/sus-02-right-sizing/) Right-sizing, the per-component version of the same problem
- [`SEC 2.1`](/architecture/pillars/security/sec-02-segmentation/#sec-21-choose-the-isolation-strength-from-the-protection-need-not-from-a-default) Isolation strength, which bounds how far density can go
- [`REL 4`](/architecture/pillars/reliability/rel-04-redundancy/) Redundancy, whose standby capacity this question tries to put to work
- [`PERF 5.3`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-53-identify-the-components-that-cannot-scale-out) Non-scaling components, which decides whether active-active is available
- [`COST 3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/) Right-sizing, which agrees with this question almost everywhere
