---
id: PERF03
pillar: performance-efficiency
title: PERF 3. How do you select and size services from measured demand?
description: Some sizing decisions are a setting and some are a migration. Knowing which is which before you choose is worth more than getting the first size right.
status: draft
services: [compute-engine, postgresql-flex, kubernetes-engine]
sidebar:
  order: 12
  label: Selection & sizing
source_url: "https://framework.stackit.cloud/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/"
source_file: "docs/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing.mdx"
---

Sizing decisions are usually made once, early, from an estimate, and then inherited for years.
That would be tolerable if they were all equally easy to revise, and they are not. Some are a
setting you change on a Tuesday. Others require migrating to a new instance.

The question is therefore two questions. What size does the demand require, and how expensive is
it to be wrong.

## Best practices

- [`PERF 3.1`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-31-size-from-measured-demand-rather-than-from-a-starting-estimate) Size from measured demand rather than from a starting estimate
- [`PERF 3.2`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-32-establish-which-sizing-decisions-are-reversible-before-you-make-them) Establish which sizing decisions are reversible before you make them
- [`PERF 3.3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-33-match-the-resource-shape-to-the-workload-shape) Match the resource shape to the workload shape
- [`PERF 3.4`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-34-know-where-the-ceiling-of-the-option-you-chose-is) Know where the ceiling of the option you chose is

---

## PERF 3.1 Size from measured demand rather than from a starting estimate

**Risk if not established:** High

The initial size is a guess, and that is unavoidable. What is avoidable is that it remains the
size two years later, when there is measured demand to size from.

Measure at the resolution where saturation actually occurs. Minute-level averages hide a
five-second burst that exhausted a connection pool, and the incident report will say the system
was at forty percent utilization. [`PERF 2.2`](/architecture/pillars/performance-efficiency/perf-02-baseline/#perf-22-measure-percentiles-rather-than-averages) makes the same point about latency.

Size for the peak that matters rather than the peak that exists. A yearly spike may be better
served by degrading under [`REL 6`](/architecture/pillars/reliability/rel-06-graceful-degradation/) than by carrying capacity for it all year, which is a decision
rather than an oversight.

Include the failure case. When a zone is lost, the remaining capacity absorbs its traffic, which
is the point [`REL 7.1`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-71-size-for-measured-peaks-and-keep-headroom-for-the-peak-you-did-not-predict) makes and the reason a system running comfortably at sixty percent per zone
has no headroom at all.

**On STACKIT.** Compute Engine offers <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server/basics/machine-types/">machine type
families</LinkChip> in
fixed variants, and a configuration cannot be adapted beyond those variants. You choose from a catalogue rather than dialling in a size, which means the useful
question is which variant your measured demand lands closest to rather than what shape you would
design.

Some families use CPU overprovisioning. Where they do, sustained performance varies with what else
is running, which makes a measurement taken at one moment a weaker predictor than it looks. For a
workload with a tight [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) target, that is worth knowing before choosing on price.

Machine types are documented per region, so what is available in `eu01` and in `eu02` is worth
confirming separately rather than assuming symmetry, particularly for a design that spans both
under [`REL 4.4`](/architecture/pillars/reliability/rel-04-redundancy/#rel-44-decide-deliberately-whether-the-workload-needs-a-second-region).

**Tradeoffs.** **Cost Optimization.** Sizing from measured demand is the same exercise as [`COST 3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/)
approached from the other side, and they usually agree. Where they disagree, it is because
performance wants headroom and cost wants utilization.

**Verify.** For your largest component, what measurement produced its current size, and when was
that taken? At what resolution was the peak measured?

---

## PERF 3.2 Establish which sizing decisions are reversible before you make them

**Risk if not established:** High

Treating all sizing as adjustable leads to under-provisioning the decisions that are not, on the
reasonable assumption that they can be corrected later.

Sort them before choosing. A **setting** can be changed with a restart or less. A **replacement**
means creating a new resource and moving traffic, which is disruptive but routine. A **migration**
means moving data, which is a project with its own risk and downtime.

The ones that hurt are the ones that look like settings and are migrations. Storage that can grow
but not shrink. A partitioning scheme chosen at creation. A database performance tier that
requires a new instance.

Where a decision is expensive to reverse, the correct response is not always to over-provision. It
may be to design so the decision matters less: a component that can be replaced rather than
resized, or state held somewhere that scales independently.

**On STACKIT.** The clearest example is <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/">PostgreSQL Flex performance
classes</LinkChip>.
A performance class that turns out to be too small requires cloning the instance to a larger one. That makes the I/O tier a migration rather than a setting, and it
belongs in the design conversation rather than being discovered when the instance is under load.

The practical consequence is to allow more headroom on that particular axis than you would on one
that is adjustable, and to know the clone-and-repoint procedure before you need it. That procedure
is the same one [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it) asks you to rehearse, which is a rare case of one exercise serving two
purposes.

Compute Engine sits at the other end: a machine type change is a replacement of the instance
rather than a data migration, which is disruptive and bounded. In <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/getting-started/node-pools/">Kubernetes
Engine</LinkChip>
node pool sizes are adjustable, while the zone set of an existing pool is not, per [`REL 4.1`](/architecture/pillars/reliability/rel-04-redundancy/#rel-41-distribute-compute-across-availability-zones-according-to-the-flows-target).

**Tradeoffs.** **Cost Optimization.** Extra headroom on irreversible axes costs money continuously
to avoid a migration that may never be needed. The size of that trade depends on how confident the
demand estimate is.

**Verify.** For each sizing decision in your workload, is changing it a setting, a replacement or
a migration? Which of the migrations did you size conservatively because of that?

---

## PERF 3.3 Match the resource shape to the workload shape

**Risk if not established:** Medium

Workloads are not uniformly hungry. Some are limited by CPU, some by memory, some by I/O, some by
network. Sizing by a single dimension means over-provisioning everything else to get enough of the
one that binds.

Establish which resource actually limits the workload before choosing, which is the same
measurement [`PERF 7.1`](/architecture/pillars/performance-efficiency/perf-07-evidence-based-optimization/#perf-71-profile-to-find-the-constraint-before-changing-anything) needs. A memory-bound service on a CPU-optimized instance wastes cores to
obtain RAM, and the reverse wastes RAM to obtain cores.

Watch for workloads that change shape. A service that was CPU-bound before a caching layer was
added may be memory-bound after it, and the instance chosen for the old shape is now wrong in a
different direction.

I/O is the dimension most often forgotten, because it is not visible in the vCPU and RAM figures
that dominate the choice. A database that fits comfortably in memory and saturates its storage
tier is limited by a number nobody looked at.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server/basics/machine-types/">machine type
families</LinkChip> are
organized precisely on this axis, by vCPU-to-RAM ratio: from CPU-weighted through general purpose
to memory-weighted, plus a GPU family. Choosing a family is choosing a shape, and it is a more
consequential decision than choosing a size within one.

> From the STACKIT docs: [Machine types - EU01 › Machine type names](https://docs.stackit.cloud/products/compute-engine/server/basics/machine-types/#machine-type-names) (Source updated 24.07.2026, copied 05.10.2026)

Example of a machine type name: `c1a.8d`

The first part of the name, e.g. “**c**1a.8d”, is the variant of a machine type. Currently, there are the following variants:

| Variant | Description |
| --- | --- |
| t | Smaller instances with smaller CPU and less RAM |
| s | Processor-optimized instances with a CPU/RAM-ratio of 1:1 |
| c | Processor-optimized instances with a CPU/RAM-ratio of 1:2 |
| g | General instances with a CPU/RAM-ratio of 1:4 |
| m | Memory-optimized instances with a CPU/RAM-ratio of 1:8 |
| b | Large, memory-optimized instances with CPU/RAM-ratio of 1:16 or higher |
| n | Instances with NVIDIA GPUs |

Two family-level properties belong in the choice rather than being discovered later. Some families
use CPU overprovisioning, and machine types providing AMD SEV or NVIDIA GPUs cannot be live
migrated, which means maintenance on those is disruptive rather than transparent. The second point
is a [`REL 3.1`](/architecture/pillars/reliability/rel-03-failure-mode-analysis/#rel-31-enumerate-how-each-component-on-a-critical-flow-can-fail) concern as much as a performance one, and it is narrower than it first appears:
it follows the confidential computing and GPU capabilities rather than the vendor.

> From the STACKIT docs: [Maintenance › Machine types with maintenance](https://docs.stackit.cloud/products/compute-engine/server/basics/maintenance/#machine-types-with-maintenance) (Source updated 06.07.2026, copied 05.10.2026)

Maintenance windows are scheduled for these machine types.

| Machine type variants | Description |
| --- | --- |
| m1a.*cd | Instances providing AMD SEV. |
| n1.*d.g* | Instances with NVIDIA A100 80 GB Tensor Core GPUs. |
| n2.*d.g* | Instances with NVIDIA L40S 48 GB GPUs. |
| n3.*d.g* | Instances with NVIDIA H100 GPUs. |

For managed databases, the shape choice appears as
<LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/">flavors</LinkChip>
in compute-optimized, memory-optimized and processor-optimized variants, with the I/O dimension
handled separately as the performance class from [`PERF 3.2`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/#perf-32-establish-which-sizing-decisions-are-reversible-before-you-make-them). Those two are chosen independently,
which is useful and also means both can be wrong independently.

> From the STACKIT docs: [Flavors and performance classes › Flavors](https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/#flavors) (Source updated 06.07.2026, copied 05.10.2026)

| Description | ID | CPU | RAM | `max_connections` | `shared_buffers` | `work_mem` | `maintenance_work_mem` | `effective_cache_size` |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Small, Compute optimized | 2.4 | 2 | 4 GB | 95 | 950 MB | 14 MB | 380 MB | 2660 MB |
| Small, Memory optimized | 2.16 | 2 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| Medium, Compute optimized | 4.8 | 4 | 8 GB | 195 | 1950 MB | 14 MB | 780 MB | 5460 MB |
| Medium, Memory optimized | 4.32 | 4 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| Large, Processor optimized | 8.16 | 8 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| X-Large, Compute optimized | 16.32 | 16 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| X-Large, Memory optimized | 16.128 | 16 | 128 GB | 3170 | 31950 MB | 14 MB | 12780 MB | 89460 MB |

### Notes

- CPU and RAM is always per node.
- The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the `max_connections` limit.

**Tradeoffs.** **Cost Optimization.** Matching the shape usually reduces cost, since it removes
the over-provisioning of the dimensions that do not bind. This is one of the places where the two
pillars agree without qualification.

**Verify.** For your largest component, which resource is the binding constraint? Does the family
you chose weight that resource, or did you choose on total size?

---

## PERF 3.4 Know where the ceiling of the option you chose is

**Risk if not established:** High

Every option has a maximum. The question is whether you find it during planning or during an
incident, and the second is considerably more expensive.

Establish the ceiling for each component on a ranked flow: the largest variant available, the
documented quota, and the point at which the architecture stops scaling regardless of the
resource. [`REL 7.3`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-73-find-every-scaling-limit-before-you-approach-it) covers the same ground from the availability side.

Compare the ceiling with projected demand, not current demand. A component at thirty percent of
its maximum with demand doubling annually has about eighteen months, and eighteen months is
roughly how long an architectural change takes to plan and execute.

Where the ceiling is close, the options are to change the architecture, to shard, or to accept a
known limit and plan for it. All three are better than reaching it unexpectedly.

**On STACKIT.** Ceilings are documented, which makes this a reading exercise rather than a
discovery exercise.

Managed database performance classes state their maximum I/O capacity per class, with the top
class an order of magnitude above the bottom, so the ceiling of a given tier is knowable before
you choose it. Beyond the largest class, the answer stops being a bigger instance and becomes a
change to the data architecture, which is [`PERF 4`](/architecture/pillars/performance-efficiency/perf-04-data-design/).

> From the STACKIT docs: [Flavors and performance classes › Performance Classes](https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/#performance-classes) (Source updated 06.07.2026, copied 05.10.2026)

| Description | ID | Max. IOPS | Max. throughput (MB/s) |
| --- | --- | --- | --- |
| Performance class 2 | `premium-perf2-stackit` | 1000 | 100 |
| Performance class 4 | `premium-perf4-stackit` | 2000 | 150 |
| Performance class 6 | `premium-perf6-stackit` | 5000 | 200 |
| Performance class 8 | `premium-perf8-stackit` | 10000 | 250 |
| Performance class 10 | `premium-perf10-stackit` | 15000 | 300 |
| Performance class 12 | `premium-perf12-stackit` | 20000 | 350 |

Currently, we offer three types of instances. For each type there is a different set of flavors available.

<LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/basics/operations/quotas-and-limits/">Kubernetes Engine quotas and
limits</LinkChip>
documents cluster and node pool maxima, and notes that clusters in certain network configurations
have lower node limits because of address space. That is an emergent limit made visible, and it is
exactly the kind that otherwise appears at the worst moment.

<LinkChip href="https://docs.stackit.cloud/platform/resource-manager/basics/projects/">Project quotas</LinkChip>
bound consumption per project, and they cover IaaS and Cloud Foundry resources rather than
everything in a project, so they are a partial ceiling rather than
a complete one. Treat a quota error as a capacity signal rather than a transient fault, as
[`REL 5.2`](/architecture/pillars/reliability/rel-05-resilient-interactions/#rel-52-retry-with-backoff-and-jitter-and-only-what-is-safe-to-retry) notes.

**Tradeoffs.** Little beyond analysis time. The cost is discovering an architectural ceiling that
requires real work to raise, which is unwelcome and much cheaper to learn now.

**Verify.** For your critical flow, what is the binding ceiling and how far is projected demand
from it? How many months does that leave, and how long would raising it take?

---

## Related

- [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) Targets, which sizing has to satisfy
- [`PERF 4`](/architecture/pillars/performance-efficiency/perf-04-data-design/) Data design, which is where the answer lies once a bigger instance stops working
- [`PERF 5`](/architecture/pillars/performance-efficiency/perf-05-scaling/) Scaling, the alternative to sizing up
- [`REL 7`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/) Scaling and headroom, the same decisions serving survival rather than speed
- [`COST 3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/) Right-sizing, which usually agrees and occasionally does not
