---
id: SUS03
pillar: sustainability
title: SUS 3. How do you shape demand so that less capacity is needed?
description: Peak demand determines how much infrastructure exists, not average demand. Moving work off the peak reduces the hardware rather than merely its utilization.
status: draft
services: [kubernetes-engine]
sidebar:
  order: 12
  label: Demand shaping
source_url: "https://framework.stackit.cloud/architecture/pillars/sustainability/sus-03-demand-shaping/"
source_file: "docs/architecture/pillars/sustainability/sus-03-demand-shaping.mdx"
---

Capacity is provisioned for the peak. Everything below the peak is spare, so the shape of demand
determines how much infrastructure exists rather than merely how busy it is.

That makes demand shaping structurally different from right-sizing. [`SUS 2`](/architecture/pillars/sustainability/sus-02-right-sizing/) fits the allocation to
the peak. This question changes the peak, which is the only move that reduces the hardware itself.

## Best practices

- [`SUS 3.1`](/architecture/pillars/sustainability/sus-03-demand-shaping/#sus-31-establish-which-work-has-a-real-deadline-and-which-merely-has-a-schedule) Establish which work has a real deadline and which merely has a schedule
- [`SUS 3.2`](/architecture/pillars/sustainability/sus-03-demand-shaping/#sus-32-batch-and-defer-what-does-not-need-to-happen-now) Batch and defer what does not need to happen now
- [`SUS 3.3`](/architecture/pillars/sustainability/sus-03-demand-shaping/#sus-33-flatten-the-peak-because-the-peak-is-what-sizes-the-estate) Flatten the peak, because the peak is what sizes the estate
- [`SUS 3.4`](/architecture/pillars/sustainability/sus-03-demand-shaping/#sus-34-ask-whether-the-work-needs-to-happen-at-all) Ask whether the work needs to happen at all

---

## SUS 3.1 Establish which work has a real deadline and which merely has a schedule

**Risk if not established:** Medium

A great deal of scheduled work runs at a particular time because somebody chose that time once,
not because anything depends on it. That distinction is invisible until somebody asks.

Go through the scheduled and background work and ask, per item, what breaks if it runs an hour
later, or overnight, or on any of the next three days. The answers separate into deadlines that
are real, deadlines that are conventions, and work with no deadline at all.

The recurring finds: reports generated before anyone reads them, indexing jobs that could run
during the quiet period, data exports scheduled at the top of the hour alongside everything else,
and cleanup tasks running during peak because that is when they were first written.

Watch for accidental synchronization. Work scheduled on round times converges: a dozen jobs at
midnight produce a peak that a dozen jobs spread across the night would not.

**On STACKIT.** Scheduled execution is available through
<LinkChip href="https://docs.stackit.cloud/products/integration/automation-service/">Automation Service</LinkChip> for
recurring platform work and through whatever scheduler your runtime provides for application work.
Neither decides which work has a deadline, which is the part of this best practice that is yours.

**Tradeoffs.** **Operational Excellence.** Analysis time, and the answers are frequently that the
deadline is real. The value is in the subset where it is not, which is usually larger than expected.

**Verify.** List your scheduled jobs and the time each runs. For how many can you state what
breaks if it ran three hours later?

---

## SUS 3.2 Batch and defer what does not need to happen now

**Risk if not established:** Medium

Work that arrives continuously and is processed immediately requires capacity sized for its
arrival rate. The same work batched requires capacity sized for the batch, which can run when
there is capacity to spare.

Deferral moves work off the peak. Batching reduces the per-item overhead, which is [`PERF 6.3`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-63-batch-what-is-chatty-and-compress-what-travels) from
the performance side and produces less total work rather than merely rescheduling it.

Both change the failure behaviour, which is the cost. Deferred work needs somewhere to wait, and
that queue is a component with its own availability and its own failure modes, as [`REL 6.1`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-61-decide-in-advance-which-functions-may-be-dropped-and-in-what-order)
describes. A batch that fails loses more than a single item would.

Bound the deferral. Work that can be postponed indefinitely tends to be, and a backlog that grows
faster than it drains is a capacity problem rather than a saving.

**On STACKIT.** Deferral needs durable storage for the pending work.
<LinkChip href="https://docs.stackit.cloud/products/messaging/rabbitmq/">RabbitMQ</LinkChip> is the managed queue and
<LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/">Object Storage</LinkChip> suits deferred bulk
work. Each adds a component that has to be as available as the deferral strategy assumes, which
belongs in the composition from [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on) rather than being treated as free.

**Tradeoffs.** **Performance Efficiency.** Batching adds latency for the first item in a batch,
which makes it right for throughput-oriented work and wrong for interactive requests.
**Reliability.** A larger unit of work loses more when it fails.

**Verify.** What proportion of your compute happens in response to an immediate request, and what
proportion is background work? Of the background work, how much runs during your peak?

---

## SUS 3.3 Flatten the peak, because the peak is what sizes the estate

**Risk if not established:** Medium

The ratio between peak and average demand is the multiplier on your infrastructure. A workload
with a peak three times its average needs three times the capacity that its average work requires,
and the difference is idle for most of the day.

Reducing that ratio is worth more than any efficiency improvement at the same scale, because it
removes hardware rather than making hardware busier.

The levers, in roughly the order they pay: move deferrable work away from the peak under `SUS
3.2`, spread scheduled work rather than converging it on round times, shed or delay low-value
traffic under [`REL 6.3`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-63-shed-load-at-the-edge-before-the-system-saturates), and where the peak is genuinely inherent, decide whether serving it is
worth the standing capacity or whether degrading is acceptable.

Distinguish peaks you can move from peaks you cannot. User traffic follows people and is largely
fixed. Batch processing, replication, backups and reporting are yours to schedule, and they
frequently coincide with the user peak for no reason other than convention.

**On STACKIT.** Where scaling is fast enough, a peak can be met with capacity that exists only
during it, which converts a standing allocation into a temporary one. Whether that works depends
on the provisioning time from [`PERF 5.4`](/architecture/pillars/performance-efficiency/perf-05-scaling/#perf-54-account-for-what-scaling-costs-in-time-and-in-money), which is workload-specific and worth measuring.

Backup windows are the schedulable peak most often left where it was set. The scheduling in
<LinkChip href="https://docs.stackit.cloud/products/compute-engine/server-backup-management/getting-started/configure-a-server-backup-management-service/">Server Backup
Management</LinkChip>
and in the managed databases is configurable, and moving it away from the user peak costs nothing.

**Tradeoffs.** **Reliability.** A flatter profile leaves less absorbed slack for an unexpected
spike, which is [`REL 7.1`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-71-size-for-measured-peaks-and-keep-headroom-for-the-peak-you-did-not-predict) and has to be sized deliberately rather than emerging.

**Verify.** What is the ratio between your peak and average demand? Which components of the peak
are schedulable rather than user-driven?

---

## SUS 3.4 Ask whether the work needs to happen at all

**Risk if not established:** Medium

The strongest form of demand shaping is removing the demand. Nothing in a normal development
process asks this question, so it is asked only when somebody decides to.

The recurring finds are consistent across estates: reports produced on a schedule that nobody has
opened in a year, data pipelines feeding a dashboard that was replaced, jobs that were part of a
migration that finished, and recomputation of values that have not changed.

The check is whether the output is consumed. That is answerable with instrumentation under `OPS
7.3`, and it is frequently not instrumented precisely because nobody suspected the answer would be
interesting.

This overlaps with [`SUS 7`](/architecture/pillars/sustainability/sus-07-efficiency-at-the-top/) and with [`SUS 6`](/architecture/pillars/sustainability/sus-06-shut-down-idle/), and the overlap is the point: work nobody consumes,
capacity nobody uses and data nobody reads are the same finding arriving through three different
questions. Whichever question surfaces it, the action is the same.

**On STACKIT.** No platform feature identifies work nobody needs. Where the output lands in
<LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/">Object Storage</LinkChip> or a database,
access patterns are observable, and an object nobody has read since it was written is a strong
signal about the job that produced it.

**Tradeoffs.** **Operational Excellence.** Removing a scheduled job requires establishing that it
is not needed, which is work and carries a small chance of being wrong. That asymmetry is why
accumulated jobs are never removed by default.

**Verify.** For your scheduled jobs, which produced output that anybody read in the last quarter?
How would you find out?

---

## Related

- [`SUS 4`](/architecture/pillars/sustainability/sus-04-utilization-density/) Utilization density, which is what a flatter profile enables
- [`SUS 6`](/architecture/pillars/sustainability/sus-06-shut-down-idle/) Shutting down idle, the same question applied to capacity rather than work
- [`REL 6.3`](/architecture/pillars/reliability/rel-06-graceful-degradation/#rel-63-shed-load-at-the-edge-before-the-system-saturates) Load shedding, the deliberate version of not serving the peak
- [`PERF 6.3`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-63-batch-what-is-chatty-and-compress-what-travels) Batching, the same technique serving throughput
- [`REL 7.1`](/architecture/pillars/reliability/rel-07-scaling-and-headroom/#rel-71-size-for-measured-peaks-and-keep-headroom-for-the-peak-you-did-not-predict) Headroom, which the peak determines
