PERF 3. How do you select and size services from measured demand?
Last updated on
Sizing decisions are usually made once, early, from an estimate, and then inherited for years. That would be tolerable if they were all equally easy to revise, and they are not. Some are a setting you change on a Tuesday. Others require migrating to a new instance.
The question is therefore two questions. What size does the demand require, and how expensive is it to be wrong.
Best practices
Section titled “Best practices”PERF 3.1Size from measured demand rather than from a starting estimatePERF 3.2Establish which sizing decisions are reversible before you make themPERF 3.3Match the resource shape to the workload shapePERF 3.4Know where the ceiling of the option you chose is
PERF 3.1 Size from measured demand rather than from a starting estimate
Section titled “PERF 3.1 Size from measured demand rather than from a starting estimate”Risk if not established: High
The initial size is a guess, and that is unavoidable. What is avoidable is that it remains the size two years later, when there is measured demand to size from.
Measure at the resolution where saturation actually occurs. Minute-level averages hide a
five-second burst that exhausted a connection pool, and the incident report will say the system
was at forty percent utilization. PERF 2.2 makes the same point about latency.
Size for the peak that matters rather than the peak that exists. A yearly spike may be better
served by degrading under REL 6 than by carrying capacity for it all year, which is a decision
rather than an oversight.
Include the failure case. When a zone is lost, the remaining capacity absorbs its traffic, which
is the point REL 7.1 makes and the reason a system running comfortably at sixty percent per zone
has no headroom at all.
On STACKIT. Compute Engine offers machine type families in fixed variants, and a configuration cannot be adapted beyond those variants. You choose from a catalogue rather than dialling in a size, which means the useful question is which variant your measured demand lands closest to rather than what shape you would design.
Some families use CPU overprovisioning. Where they do, sustained performance varies with what else
is running, which makes a measurement taken at one moment a weaker predictor than it looks. For a
workload with a tight PERF 1 target, that is worth knowing before choosing on price.
Machine types are documented per region, so what is available in eu01 and in eu02 is worth
confirming separately rather than assuming symmetry, particularly for a design that spans both
under REL 4.4.
Tradeoffs. Cost Optimization. Sizing from measured demand is the same exercise as COST 3
approached from the other side, and they usually agree. Where they disagree, it is because
performance wants headroom and cost wants utilization.
Verify. For your largest component, what measurement produced its current size, and when was that taken? At what resolution was the peak measured?
PERF 3.2 Establish which sizing decisions are reversible before you make them
Section titled “PERF 3.2 Establish which sizing decisions are reversible before you make them”Risk if not established: High
Treating all sizing as adjustable leads to under-provisioning the decisions that are not, on the reasonable assumption that they can be corrected later.
Sort them before choosing. A setting can be changed with a restart or less. A replacement means creating a new resource and moving traffic, which is disruptive but routine. A migration means moving data, which is a project with its own risk and downtime.
The ones that hurt are the ones that look like settings and are migrations. Storage that can grow but not shrink. A partitioning scheme chosen at creation. A database performance tier that requires a new instance.
Where a decision is expensive to reverse, the correct response is not always to over-provision. It may be to design so the decision matters less: a component that can be replaced rather than resized, or state held somewhere that scales independently.
On STACKIT. The clearest example is PostgreSQL Flex performance classes . A performance class that turns out to be too small requires cloning the instance to a larger one. That makes the I/O tier a migration rather than a setting, and it belongs in the design conversation rather than being discovered when the instance is under load.
The practical consequence is to allow more headroom on that particular axis than you would on one
that is adjustable, and to know the clone-and-repoint procedure before you need it. That procedure
is the same one REL 8.3 asks you to rehearse, which is a rare case of one exercise serving two
purposes.
Compute Engine sits at the other end: a machine type change is a replacement of the instance
rather than a data migration, which is disruptive and bounded. In Kubernetes
Engine
node pool sizes are adjustable, while the zone set of an existing pool is not, per REL 4.1.
Tradeoffs. Cost Optimization. Extra headroom on irreversible axes costs money continuously to avoid a migration that may never be needed. The size of that trade depends on how confident the demand estimate is.
Verify. For each sizing decision in your workload, is changing it a setting, a replacement or a migration? Which of the migrations did you size conservatively because of that?
PERF 3.3 Match the resource shape to the workload shape
Section titled “PERF 3.3 Match the resource shape to the workload shape”Risk if not established: Medium
Workloads are not uniformly hungry. Some are limited by CPU, some by memory, some by I/O, some by network. Sizing by a single dimension means over-provisioning everything else to get enough of the one that binds.
Establish which resource actually limits the workload before choosing, which is the same
measurement PERF 7.1 needs. A memory-bound service on a CPU-optimized instance wastes cores to
obtain RAM, and the reverse wastes RAM to obtain cores.
Watch for workloads that change shape. A service that was CPU-bound before a caching layer was added may be memory-bound after it, and the instance chosen for the old shape is now wrong in a different direction.
I/O is the dimension most often forgotten, because it is not visible in the vCPU and RAM figures that dominate the choice. A database that fits comfortably in memory and saturates its storage tier is limited by a number nobody looked at.
On STACKIT. The machine type families are organized precisely on this axis, by vCPU-to-RAM ratio: from CPU-weighted through general purpose to memory-weighted, plus a GPU family. Choosing a family is choosing a shape, and it is a more consequential decision than choosing a size within one.
Example of a machine type name: c1a.8d
The first part of the name, e.g. “c1a.8d”, is the variant of a machine type. Currently, there are the following variants:
| Variant | Description |
|---|---|
| t | Smaller instances with smaller CPU and less RAM |
| s | Processor-optimized instances with a CPU/RAM-ratio of 1:1 |
| c | Processor-optimized instances with a CPU/RAM-ratio of 1:2 |
| g | General instances with a CPU/RAM-ratio of 1:4 |
| m | Memory-optimized instances with a CPU/RAM-ratio of 1:8 |
| b | Large, memory-optimized instances with CPU/RAM-ratio of 1:16 or higher |
| n | Instances with NVIDIA GPUs |
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Two family-level properties belong in the choice rather than being discovered later. Some families
use CPU overprovisioning, and machine types providing AMD SEV or NVIDIA GPUs cannot be live
migrated, which means maintenance on those is disruptive rather than transparent. The second point
is a REL 3.1 concern as much as a performance one, and it is narrower than it first appears:
it follows the confidential computing and GPU capabilities rather than the vendor.
Maintenance windows are scheduled for these machine types.
| Machine type variants | Description |
|---|---|
| m1a.*cd | Instances providing AMD SEV. |
| n1.*d.g* | Instances with NVIDIA A100 80 GB Tensor Core GPUs. |
| n2.*d.g* | Instances with NVIDIA L40S 48 GB GPUs. |
| n3.*d.g* | Instances with NVIDIA H100 GPUs. |
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
For managed databases, the shape choice appears as
flavors
in compute-optimized, memory-optimized and processor-optimized variants, with the I/O dimension
handled separately as the performance class from PERF 3.2. Those two are chosen independently,
which is useful and also means both can be wrong independently.
| Description | ID | CPU | RAM | max_connections | shared_buffers | work_mem | maintenance_work_mem | effective_cache_size |
|---|---|---|---|---|---|---|---|---|
| Small, Compute optimized | 2.4 | 2 | 4 GB | 95 | 950 MB | 14 MB | 380 MB | 2660 MB |
| Small, Memory optimized | 2.16 | 2 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| Medium, Compute optimized | 4.8 | 4 | 8 GB | 195 | 1950 MB | 14 MB | 780 MB | 5460 MB |
| Medium, Memory optimized | 4.32 | 4 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| Large, Processor optimized | 8.16 | 8 | 16 GB | 385 | 3950 MB | 14 MB | 1580 MB | 11060 MB |
| X-Large, Compute optimized | 16.32 | 16 | 32 GB | 785 | 7950 MB | 14 MB | 3180 MB | 22260 MB |
| X-Large, Memory optimized | 16.128 | 16 | 128 GB | 3170 | 31950 MB | 14 MB | 12780 MB | 89460 MB |
Notes
- CPU and RAM is always per node.
- The system uses up to 15 connections for internal essential processes such as backup, monitoring, etc. These connections will be counted towards the
max_connectionslimit.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Tradeoffs. Cost Optimization. Matching the shape usually reduces cost, since it removes the over-provisioning of the dimensions that do not bind. This is one of the places where the two pillars agree without qualification.
Verify. For your largest component, which resource is the binding constraint? Does the family you chose weight that resource, or did you choose on total size?
PERF 3.4 Know where the ceiling of the option you chose is
Section titled “PERF 3.4 Know where the ceiling of the option you chose is”Risk if not established: High
Every option has a maximum. The question is whether you find it during planning or during an incident, and the second is considerably more expensive.
Establish the ceiling for each component on a ranked flow: the largest variant available, the
documented quota, and the point at which the architecture stops scaling regardless of the
resource. REL 7.3 covers the same ground from the availability side.
Compare the ceiling with projected demand, not current demand. A component at thirty percent of its maximum with demand doubling annually has about eighteen months, and eighteen months is roughly how long an architectural change takes to plan and execute.
Where the ceiling is close, the options are to change the architecture, to shard, or to accept a known limit and plan for it. All three are better than reaching it unexpectedly.
On STACKIT. Ceilings are documented, which makes this a reading exercise rather than a discovery exercise.
Managed database performance classes state their maximum I/O capacity per class, with the top
class an order of magnitude above the bottom, so the ceiling of a given tier is knowable before
you choose it. Beyond the largest class, the answer stops being a bigger instance and becomes a
change to the data architecture, which is PERF 4.
| Description | ID | Max. IOPS | Max. throughput (MB/s) |
|---|---|---|---|
| Performance class 2 | premium-perf2-stackit | 1000 | 100 |
| Performance class 4 | premium-perf4-stackit | 2000 | 150 |
| Performance class 6 | premium-perf6-stackit | 5000 | 200 |
| Performance class 8 | premium-perf8-stackit | 10000 | 250 |
| Performance class 10 | premium-perf10-stackit | 15000 | 300 |
| Performance class 12 | premium-perf12-stackit | 20000 | 350 |
Currently, we offer three types of instances. For each type there is a different set of flavors available.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Kubernetes Engine quotas and limits documents cluster and node pool maxima, and notes that clusters in certain network configurations have lower node limits because of address space. That is an emergent limit made visible, and it is exactly the kind that otherwise appears at the worst moment.
Project quotas
bound consumption per project, and they cover IaaS and Cloud Foundry resources rather than
everything in a project, so they are a partial ceiling rather than
a complete one. Treat a quota error as a capacity signal rather than a transient fault, as
REL 5.2 notes.
Tradeoffs. Little beyond analysis time. The cost is discovering an architectural ceiling that requires real work to raise, which is unwelcome and much cheaper to learn now.
Verify. For your critical flow, what is the binding ceiling and how far is projected demand from it? How many months does that leave, and how long would raising it take?
Related
Section titled “Related”PERF 1Targets, which sizing has to satisfyPERF 4Data design, which is where the answer lies once a bigger instance stops workingPERF 5Scaling, the alternative to sizing upREL 7Scaling and headroom, the same decisions serving survival rather than speedCOST 3Right-sizing, which usually agrees and occasionally does not