---
id: COST04
pillar: cost-optimization
title: COST 4. How do you choose the service and tier that matches the requirement?
description: Tier decisions are made once and rarely revisited, even when the requirement that justified them has disappeared. Know what actually drives each bill.
status: draft
services: [object-storage, postgresql-flex]
sidebar:
  order: 13
  label: Service selection
source_url: "https://framework.stackit.cloud/architecture/pillars/cost-optimization/cost-04-service-selection/"
source_file: "docs/architecture/pillars/cost-optimization/cost-04-service-selection.mdx"
---

Two mistakes account for most of the cost in this question, and they run in opposite directions.
Choosing a service because it is familiar rather than because it fits, and choosing a tier that
meets a requirement nobody has stated.

Both are made once, early, and inherited. A tier chosen for a peak that never arrived costs the
difference every month for years, and nothing draws attention to it.

## Best practices

- [`COST 4.1`](/architecture/pillars/cost-optimization/cost-04-service-selection/#cost-41-choose-the-cheapest-option-that-meets-the-stated-requirement) Choose the cheapest option that meets the stated requirement
- [`COST 4.2`](/architecture/pillars/cost-optimization/cost-04-service-selection/#cost-42-compare-managed-against-self-operated-with-the-labour-counted) Compare managed against self-operated with the labour counted
- [`COST 4.3`](/architecture/pillars/cost-optimization/cost-04-service-selection/#cost-43-know-what-actually-drives-the-bill-for-each-option) Know what actually drives the bill for each option
- [`COST 4.4`](/architecture/pillars/cost-optimization/cost-04-service-selection/#cost-44-re-examine-when-the-requirement-changes) Re-examine when the requirement changes

---

## COST 4.1 Choose the cheapest option that meets the stated requirement

**Risk if not established:** Medium

The requirement has to exist first, which is why this question depends on [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) and [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/)
having produced numbers. Without them, tier selection is a judgement about how much reliability
and speed feels right, and that judgement is consistently generous.

Choose against the stated requirement rather than against a possible future one. Provisioning for
a requirement you might have later means paying for it from today, and the alternative is usually
available: choose for now and re-examine under [`COST 4.4`](/architecture/pillars/cost-optimization/cost-04-service-selection/#cost-44-re-examine-when-the-requirement-changes).

Where the difference between tiers is small, take the higher one and stop thinking about it. Where
it is a factor of several, it warrants the analysis. Spending an afternoon to save a few euros a
month is its own kind of waste.

Watch for tiers chosen for a property that is not actually used. A high-availability configuration
for a workload whose stated RTO is a day, or a performance class chosen for headroom on an axis
the workload never loads.

**On STACKIT.** Options are structured differently per service and the dimensions are worth
reading before choosing: machine type variants for Compute Engine, <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/reference/flavors-and-performance-classes-of-postgresql-flex/">flavors and performance
classes</LinkChip>
for managed databases, and service plans elsewhere.

The single-instance against replica-set choice for a managed database is the one with the largest
cost difference and the clearest requirement behind it. [`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) produces the target that decides
it, and the <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/basics/plan-your-postgresql-flex-instance/">planning
guidance</LinkChip>
names the three-replica type for production use, which is a recommendation rather than a
requirement your particular workload necessarily has.

**Tradeoffs.** **Reliability.** The cheapest option that meets the requirement has no margin for
the requirement being wrong, which is why [`REL 1.1`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-11-quantify-what-an-outage-and-what-data-loss-cost-the-business) insists the requirement come from the business
rather than from a preference.

**Verify.** For each service you use, which tier is selected and which stated requirement
justifies it? How many were chosen without a stated requirement?

---

## COST 4.2 Compare managed against self-operated with the labour counted

**Risk if not established:** Medium

A comparison of the invoice line alone reliably concludes that self-operated is cheaper, because
its dominant cost is people and people are not on the invoice.

Count what [`COST 1.2`](/architecture/pillars/cost-optimization/cost-01-cost-model/#cost-12-include-operational-labour-not-only-resource-consumption) describes: building it, patching it, monitoring it, backing it up, carrying
the pager, and the expertise that has to exist. Then add the failure modes you inherit, since a
self-operated database has an availability that depends on your operations rather than on a
published commitment.

The comparison usually reverses once labour is counted, and not always. Self-operating makes sense
where the managed option does not fit the requirement, where you need a version or configuration
the service does not offer, or where the scale is large enough that the per-unit premium exceeds
the labour.

Include the reversibility. Moving from managed to self-operated later is possible; the reverse
usually is too. What is harder to reverse is the expertise decision, because a team that has never
operated the component cannot start quickly.

**On STACKIT.** The choice exists at most layers. A database on Compute Engine against <LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/">PostgreSQL
Flex</LinkChip>, your own CI against
<LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/basics/stackit-pipelines/">STACKIT
Pipelines</LinkChip>,
your own backup automation against <LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/">Server Backup
Management</LinkChip>, your
own scheduler against <LinkChip href="https://docs.stackit.cloud/products/integration/automation-service/">Automation
Service</LinkChip>.

One asymmetry is worth stating: a managed service comes with a published availability commitment
in its <LinkChip href="https://stackit.com/en/gtc/service-certificates">service certificate</LinkChip>, which is an input
to the composition in [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on). A self-operated equivalent has whatever availability your
operations achieve, and that number is not published anywhere because nobody has measured it.

**Tradeoffs.** **Sovereignty & Compliance.** Managed means the provider operates it, which changes
the shared responsibility split that [`SOV 8`](/architecture/pillars/sovereignty/sov-08-compliance-mapping/) maps. That is a factor rather than an objection and
it belongs in the comparison.

**Verify.** For each self-operated component, what would the managed equivalent cost, and how many
engineer-hours per month does the current arrangement consume?

---

## COST 4.3 Know what actually drives the bill for each option

**Risk if not established:** Medium

Cost intuition transfers badly between services. The dimension that dominates one is negligible in
another, and a design optimized for the wrong dimension saves nothing.

For each service on a ranked flow, establish which dimension carries the cost. Sometimes it is
capacity, sometimes throughput, sometimes the number of operations, sometimes data transfer,
sometimes the number of instances regardless of their size.

The ones that surprise people are the operation-count and transfer dimensions, because they scale
with behaviour rather than with size. A storage bucket holding very little data and receiving
millions of small requests can cost more than one holding a great deal and receiving few.

Knowing the dominant dimension is what makes an optimization worth doing. Compressing objects
reduces a capacity bill and does nothing for a request-count bill; batching does the reverse.
[`PERF 6.3`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-63-batch-what-is-chatty-and-compress-what-travels) describes the same techniques from the performance side and they pay differently here.

**On STACKIT.** Object Storage is the clearest case where behaviour rather than volume drives both
cost and performance. The <LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/tutorials/optimize-object-storage-performance/">performance
guidance</LinkChip>
covers object size, request rate and parallelism, and those are the same dimensions to think about
when the bill is the concern.

Inter-zone traffic is the dimension most often forgotten in a multi-zone design, and it is a
direct consequence of the redundancy chosen under [`REL 4.1`](/architecture/pillars/reliability/rel-04-redundancy/#rel-41-distribute-compute-across-availability-zones-according-to-the-flows-target). A chatty component distributed across
zones pays for every call, which is a cost that did not exist before the topology changed.

**Tradeoffs.** Little. This is an analysis that changes which optimizations are worth attempting,
and getting it wrong means effort spent on the dimension that was not the bill.

**Verify.** For your three largest cost items, which dimension drives each? Which of those did you
have to look up rather than know?

---

## COST 4.4 Re-examine when the requirement changes

**Risk if not established:** Medium

Tier decisions are among the stickiest in an estate. They are made at design time, they work, and
nothing about a working configuration draws attention to itself.

Requirements change underneath them. A workload that was customer-facing becomes internal. A
reporting database that needed to be fast now runs overnight. A service that carried regulated
data no longer does. Each of those changes the tier that is justified, and none of them triggers a
review.

Tie the re-examination to the change rather than to a calendar. A change in classification under
[`SEC 3`](/architecture/pillars/security/sec-03-data-classification/), a change in the flow ranking under [`REL 2.2`](/architecture/pillars/reliability/rel-02-critical-flows/#rel-22-rank-flows-by-the-consequence-of-failure-rather-than-by-traffic-volume), or a change in the stated target under
[`REL 1`](/architecture/pillars/reliability/rel-01-reliability-targets/) and [`PERF 1`](/architecture/pillars/performance-efficiency/perf-01-performance-targets/) should each prompt the question of whether the tier still fits.

Look in both directions. The reflex is to check whether a tier is now too small; the saving is
usually in the ones that are now too large.

**On STACKIT.** Whether the re-examination can act on its conclusion depends on the reversibility
sort from [`COST 3.3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/#cost-33-know-which-resizings-are-cheap-and-which-are-migrations). A service plan that can be changed is worth reviewing frequently; one that
requires a migration is worth reviewing when something else already justifies the disruption, such
as a version upgrade under [`OPS 6.3`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-63-keep-platform-versions-aligned-and-upgrade-in-the-same-order-as-you-deploy).

**Tradeoffs.** **Operational Excellence.** Acting on a re-examination means a change with its own
risk, which is why the reversibility sort decides whether it is worth acting on rather than
recording.

**Verify.** Which of your tier decisions were made more than a year ago? For each, has the
requirement that justified it changed?

---

## Related

- [`COST 1.2`](/architecture/pillars/cost-optimization/cost-01-cost-model/#cost-12-include-operational-labour-not-only-resource-consumption) Cost model, where the labour comparison belongs
- [`COST 3.3`](/architecture/pillars/cost-optimization/cost-03-right-sizing/#cost-33-know-which-resizings-are-cheap-and-which-are-migrations) Reversibility, which decides whether a re-examination can act
- [`PERF 3`](/architecture/pillars/performance-efficiency/perf-03-service-selection-and-sizing/) Selection and sizing, the same decisions serving the performance target
- [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on) Availability composition, which service certificates feed
- [`SUS 2`](/architecture/pillars/sustainability/sus-02-right-sizing/) Right-sizing, which reaches similar conclusions from consumption
