Zum Inhalt springen
Beta

SUS 4. How do you increase utilization density?

Zuletzt aktualisiert am

Where the provider layer is already efficient, as SUS 1.3 establishes for STACKIT, the variable left to you is how much infrastructure you occupy and how much useful work happens per unit of it.

The reason density dominates is that consumption is not proportional to utilization. A lightly used server draws substantial power, occupies physical space, and represents hardware whose manufacturing footprint was paid regardless of what it subsequently does. Consolidating the same work onto fewer, better-used resources reduces consumption far more than making an individual workload marginally more efficient.

  • SUS 4.1 Consolidate onto fewer, better-utilized resources
  • SUS 4.2 Prefer active-active over idle standby
  • SUS 4.3 Account for what the platform reserves before it reaches your workload
  • SUS 4.4 Know where density conflicts with isolation, and decide rather than default

SUS 4.1 Consolidate onto fewer, better-utilized resources

Section titled “SUS 4.1 Consolidate onto fewer, better-utilized resources”

Risk if not established: Medium

Estates fragment over time. Each workload gets its own instance because that was simplest at the moment it was created, and the result is many resources each doing a little.

Consolidation is the mechanical answer: the same work on fewer, larger, better-used resources. It reduces the number of things drawing power, the number of things carrying a system overhead, and the number of things to operate, which is OPS 10 benefiting from the same change.

Two properties make a workload a consolidation candidate. It tolerates neighbours, meaning it does not need dedicated hardware for isolation or for predictable performance. And its demand profile complements others, so peaks do not coincide.

The second is the one that decides whether consolidation helps. Two workloads peaking at the same hour need the sum of their peaks whether they share a host or not; two peaking at different times need only the larger.

On STACKIT. Container platforms are the practical consolidation mechanism. In Kubernetes Engine , many workloads share a node pool, and how densely depends on the resource requests you set: requests that are generous relative to actual use reserve capacity that nothing consumes, which reproduces the fragmentation problem inside the cluster.

That makes container resource requests the same decision as instance sizing in SUS 2, applied at a finer grain and with the same tendency to round up.

Tradeoffs. Reliability. Consolidation concentrates blast radius: one host or one cluster failing affects more workloads. Performance Efficiency. Neighbours contend for shared resources, which makes performance less predictable.

Verify. How many separate compute resources does your estate run, and what is the average utilization across them? How many could be combined?


SUS 4.2 Prefer active-active over idle standby

Section titled “SUS 4.2 Prefer active-active over idle standby”

Risk if not established: Medium

This is the sharpest conflict between this pillar and Reliability, and it does not resolve in this pillar’s favour. Redundancy is deliberately reserved idle capacity, and REL 4 is not waived for environmental reasons.

What is available is the choice of shape. An active-passive design keeps a standby doing nothing until a failure that may never occur. An active-active design uses the same total capacity to serve traffic, so the redundancy is a property of how the load is spread rather than a reservation.

Where the workload allows it, active-active gives the same failure tolerance with the capacity doing useful work. Where it does not, because the component cannot run as multiple active instances under PERF 5.3, standby is the correct answer and the idle capacity is justified.

Where standby is unavoidable, two reductions remain. Keep it minimal and scale it on failover rather than mirroring production continuously, and check that the standby is actually needed at the size it is, which is SUS 2.3 distinguishing headroom from slack.

On STACKIT. The topology choice appears directly in the availability figures from REL 1.3, where a system group spanning two zones is committed to a higher availability than a single instance. What that comparison does not say is whether both instances serve traffic, which is your design decision and the one this best practice is about.

For managed databases the replica topology is fixed by the service, so the shape is not yours to choose, and the replicas do not serve reads. That capacity waits rather than works, which is the fact PERF 5.3 starts from.

Tradeoffs. Reliability. Active-active is harder to reason about, requires the workload to tolerate concurrent instances, and has failure modes that active-passive does not. It is the right default where it fits and not a universal improvement.

Verify. For each redundant component, is the redundant capacity serving traffic or waiting? For those waiting, what prevents them from serving?


SUS 4.3 Account for what the platform reserves before it reaches your workload

Section titled “SUS 4.3 Account for what the platform reserves before it reaches your workload”

Risk if not established: Medium

Allocated capacity and usable capacity are not the same number. Every layer between the hardware and your process takes something: the hypervisor, the operating system, the orchestrator, the agents.

Sizing against the nominal figure means the workload has less than intended, which produces either a performance problem or a compensating over-allocation. Both are worse than knowing the reservation.

The reservation is frequently not proportional. Where a fixed overhead exists per unit, many small units carry more total overhead than fewer large ones for the same nominal capacity, which is an argument for larger units that runs alongside the consolidation argument in SUS 4.1.

On STACKIT. Kubernetes Engine documents this precisely, and it is the clearest example of the non-proportionality. The quotas and limits page describes system resource reservations on every node, tiered so that a larger share of a small node is reserved than of a large one.

The consequence is direct: a node pool of many small nodes loses proportionally more capacity to system overhead than a pool of fewer large nodes providing the same nominal total. That is a density decision hiding inside a node size decision, and it is not visible from the machine type figures.

The same page notes that clusters in certain network configurations have lower node maxima because of address space, which is a different constraint and belongs in PERF 3.4.

Tradeoffs. Reliability. Fewer, larger nodes means a larger blast radius per node failure and coarser granularity when scaling, which is the same trade SUS 4.1 makes.

Verify. For your cluster, what is the difference between nominal and allocatable capacity? How does that proportion change with node size?


SUS 4.4 Know where density conflicts with isolation, and decide rather than default

Section titled “SUS 4.4 Know where density conflicts with isolation, and decide rather than default”

Risk if not established: Medium

Density and isolation pull against each other directly. Every boundary that SEC 2.1 describes prevents consolidation across it, and every consolidation removes a boundary.

That is a real tradeoff rather than a problem to solve. A workload whose classification requires a separate project or a dedicated instance is not a density candidate, and pursuing density there would be optimizing the wrong thing.

What this best practice asks for is that the boundary be a decision rather than a default. A workload running alone because its classification requires it is correct. A workload running alone because it was created that way is a consolidation candidate that nobody examined.

Use the classification from SEC 3 and the tier from SOV 1 to sort them. Where the protection need permits shared infrastructure with logical separation, the mechanisms in SEC 2.1 provide it and the density is available. Where it does not, the capacity is justified.

On STACKIT. The in-project separation mechanisms from SEC 2.1 are what make density and isolation partly compatible: namespaces with Kubernetes RBAC and network policy, KMS key rings, per-service access models. Each allows workloads to share infrastructure while remaining separated, with the caveat recorded there about who can bypass those boundaries.

Where the requirement is hardware-level isolation, for example under a classification that demands it, protecting data in use carries a consumption premium as well as a cost one, per SOV 5. That is a justified premium rather than waste, and worth recording as such so a later density review does not treat it as a target.

Tradeoffs. Security and Sovereignty & Compliance, both directly. This best practice exists to make the tradeoff explicit rather than to resolve it in either direction.

Verify. For each workload running on dedicated infrastructure, what requires that? How many are isolated by decision rather than by history?


  • SUS 2 Right-sizing, the per-component version of the same problem
  • SEC 2.1 Isolation strength, which bounds how far density can go
  • REL 4 Redundancy, whose standby capacity this question tries to put to work
  • PERF 5.3 Non-scaling components, which decides whether active-active is available
  • COST 3 Right-sizing, which agrees with this question almost everywhere