---
id: SEC03
pillar: security
title: SEC 3. How do you classify data, and how does the classification drive your controls?
description: Uniform controls across mixed data are too expensive for most of it and too weak where it counts. Classification makes every later control decision answerable.
status: draft
services: [object-storage, postgresql-flex]
sidebar:
  order: 12
  label: Data classification
source_url: "https://framework.stackit.cloud/architecture/pillars/security/sec-03-data-classification/"
source_file: "docs/architecture/pillars/security/sec-03-data-classification.mdx"
---

Almost everything else in this pillar depends on knowing what you are protecting. Encryption, key
ownership, access control, retention, network exposure and placement all have different right
answers for a public product catalogue and for health records, and a system that treats them the
same is wrong for both.

Teams that skip this end up applying one level of control everywhere. That is expensive for the
majority of the data and insufficient for the part that mattered, which is the worst available
combination.

## Best practices

- [`SEC 3.1`](/architecture/pillars/security/sec-03-data-classification/#sec-31-classify-every-data-set-before-deciding-where-it-lives) Classify every data set before deciding where it lives
- [`SEC 3.2`](/architecture/pillars/security/sec-03-data-classification/#sec-32-let-the-classification-drive-the-controls-mechanically) Let the classification drive the controls, mechanically
- [`SEC 3.3`](/architecture/pillars/security/sec-03-data-classification/#sec-33-find-the-copies-telemetry-backups-caches-test-data-and-exports) Find the copies: telemetry, backups, caches, test data and exports
- [`SEC 3.4`](/architecture/pillars/security/sec-03-data-classification/#sec-34-minimize-what-you-hold) Minimize what you hold

---

## SEC 3.1 Classify every data set before deciding where it lives

**Risk if not established:** Medium

Keep the scheme small. Three or four levels are enough, and more produces debates about which
category something falls into rather than decisions about how to protect it.

Classify by consequence of disclosure, not by which system holds it. The same database can contain
data at two levels, and a classification applied per system rather than per data set will protect
the wrong thing.

Two dimensions are worth separating, because they diverge. **Sensitivity** is how bad disclosure
would be, which is what this question addresses. **Regulatory exposure** is which legal regime
governs it, which is [`SOV 1`](/architecture/pillars/sovereignty/sov-01-sovereignty-tier/) and answers a different question. A data set can be commercially
sensitive without being regulated, and regulated without being especially sensitive.

Assign an owner per data set who can answer questions about it. Classification decays when the
person who made the call has left and nobody knows why the answer was what it was.

**On STACKIT.** No platform feature classifies your data. It is a business exercise and no
provider can do it for you.

Where the platform intersects is that the classification becomes an input to platform decisions:
which encryption and key ownership under [`SEC 7`](/architecture/pillars/security/sec-07-encryption/) and [`SOV 4`](/architecture/pillars/sovereignty/sov-04-key-ownership/), which placement under [`SOV 2`](/architecture/pillars/sovereignty/sov-02-placement-and-residency/), and
which retention under [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/).

For the regulatory dimension, the Advisory Framework's three sovereignty tiers are the
classification to use rather than inventing one, which is [`SOV 1`](/architecture/pillars/sovereignty/sov-01-sovereignty-tier/). Keeping the security
classification and the sovereignty tier as two related but distinct labels is more useful than
merging them, because they drive different controls.

**Tradeoffs.** None between pillars. The cost is time from people who know the business rather than
the system. The main risk is a scheme so detailed that classification becomes a project instead of a
step.

**Verify.** List your data sets and their classification. Who assigned each, and can they say why?
Which data sets are not on the list?

---

## SEC 3.2 Let the classification drive the controls, mechanically

**Risk if not established:** Medium

A classification that does not change anything is a label. The value comes from a stated mapping
from level to required controls, so that classifying a new data set determines its treatment
rather than starting a discussion.

Write the mapping once, covering at least: encryption and who holds the keys, who may access it
and under what conditions, where it may be stored and processed, how long it is kept, whether it
may appear in logs, and whether it may be copied to non-production.

The mapping is what makes the controls checkable, which is what [`SEC 1.1`](/architecture/pillars/security/sec-01-security-baseline/#sec-11-derive-the-baseline-from-the-frameworks-you-are-actually-assessed-against) needs for a baseline. It
also makes the cost defensible, since the expensive controls are visibly attached to the data that
requires them rather than applied uniformly out of caution.

Handle the mixed case explicitly. A system holding data at several levels is governed by the
highest, unless the levels are genuinely separated inside it. That is usually an argument for
separating them, since the alternative is applying the strictest treatment to everything.

**On STACKIT.** The controls the mapping refers to are the platform features covered elsewhere:
<LinkChip href="https://docs.stackit.cloud/products/security/kms/">KMS</LinkChip> for key ownership under [`SEC 7`](/architecture/pillars/security/sec-07-encryption/), <LinkChip href="https://docs.stackit.cloud/platform/access-and-identity/roles-permissions/roles-permissions/">roles
and
permissions</LinkChip>
for access under [`SEC 5`](/architecture/pillars/security/sec-05-least-privilege/), and placement under [`SOV 2`](/architecture/pillars/sovereignty/sov-02-placement-and-residency/).

<LinkChip href="https://docs.stackit.cloud/products/security/cspm/">CSPM</LinkChip> checks configuration against benchmarks
rather than against your classification, so it will find a publicly readable bucket regardless of
what is in it. Connecting a finding to the classification of the data involved is your annotation,
usually through labelling, and it is what turns a generic finding into a prioritized one.

**Tradeoffs.** **Cost Optimization.** Tiered controls are more work to design than uniform ones
and substantially cheaper to run, since the expensive treatment applies where it is warranted.

**Verify.** For your highest classification level, what does the mapping require? Pick a data set
at that level and check each requirement against reality.

---

## SEC 3.3 Find the copies: telemetry, backups, caches, test data and exports

**Risk if not established:** High

Classification is usually applied to the primary store and to nothing else. The copies are where
the exposure concentrates, because they live in places with weaker controls and broader access.

The five that recur:

**Telemetry.** Logs, traces and error reports routinely contain the data they describe. A debug
log that dumps a request body has moved regulated data into a store queried by everyone on call.
This is [`OPS 7.2`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-72-propagate-a-correlation-identifier-through-every-hop) and [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/).

**Backups.** A full copy by definition, frequently held longer than the original and with
different access. [`REL 8.5`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-85-protect-backups-to-the-classification-of-their-contents) covers the protection side.

**Caches.** A cache holds a copy of whatever it serves, with the same classification and usually
a broader read path. [`PERF 6.2`](/architecture/pillars/performance-efficiency/perf-06-reduce-work/#perf-62-cache-what-is-expensive-and-stable-and-decide-the-staleness-deliberately) adds them for latency; each one is a store to count here.

**Test data.** A copy of production in an environment with weaker controls, which is [`OPS 6.4`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-64-test-with-production-like-data-without-using-production-data).

**Exports.** The CSV someone produced for an analysis, the extract for a migration in 2023, the
report emailed to a stakeholder. These are outside every control you built and are the hardest to
find.

Trace them from the data set rather than looking for them. Ask where this data goes, not where
sensitive data is, because the second question has no complete answer.

**On STACKIT.** Telemetry destinations are explicit configuration:
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip>,
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/logs/">Logs</LinkChip>, and export through
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/">Telemetry Router</LinkChip>.
Each is a place a copy can land, and each is a placement decision under [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/) rather than an
operational detail.

<LinkChip href="https://docs.stackit.cloud/products/databases/postgresql-flex/how-tos/backup-and-clone-postgresql-flex/">Instance
cloning</LinkChip>
makes copies easy, which is useful for the rehearsals in [`REL 8.3`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-83-restore-on-a-cadence-into-a-clean-environment-and-time-it) and is exactly the mechanism to
watch here. A clone of a production database is production data with production classification.

**Tradeoffs.** **Operational Excellence.** Restricting what may appear in logs constrains
debugging, and the constraint is felt during incidents. Structured logging with explicit field
selection is the usual resolution, and it is work.

**Verify.** For your most sensitive data set, list every copy: telemetry, backups, replicas,
caches, test environments, exports. Which of those did you have to think about rather than look
up?

---

## SEC 3.4 Minimize what you hold

**Risk if not established:** Medium

Data you do not hold cannot be breached, cannot be mishandled, and needs no controls. Minimization
is the only measure in this pillar with no operational cost after the fact.

Three questions per data set. Do you need to collect it at all, or was it collected because it was
available? Do you need it in this form, or would a hash, a token or an aggregate serve? Do you
need it for this long, or is the retention a default nobody chose?

The third is where most of the reduction is available, and it is the same lever [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/) pulls for
a different reason. Data accumulates by default and is almost never deleted by default, because
nobody is ever blamed for keeping something.

Pseudonymization and tokenization reduce exposure without losing utility for many purposes. They
are not anonymization and should not be described as such, since re-identification is frequently
possible and the regulatory treatment differs.

**On STACKIT.** Retention is configured per service:
<LinkChip href="https://docs.stackit.cloud/products/storage/object-storage/">Object Storage lifecycle</LinkChip>, database
backup retention under [`REL 8.1`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-81-derive-the-backup-schedule-from-the-rpo-per-data-set), and telemetry retention under [`OPS 7.4`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-74-set-retention-by-the-value-of-each-signal-rather-than-uniformly).

The <LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> matters here: it is retained for
90 days in the Portal by default, and longer retention is a deliberate
export through Telemetry Router rather than something that accumulates silently. That is the shape
minimization wants, with retention as a decision rather than a default.

**Tradeoffs.** **Performance Efficiency.** Data removed is data unavailable for analysis, and the
value of historical data is frequently discovered after it has been deleted.
**Sovereignty & Compliance.** Regulatory retention is a floor that minimization cannot go below,
which [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) sets.

**Verify.** For your largest data set, what is the retention and who chose it? What proportion of
it has been read in the last year?

---

## Related

- [`SOV 1`](/architecture/pillars/sovereignty/sov-01-sovereignty-tier/) Sovereignty tier, the regulatory dimension of the same exercise
- [`SEC 7`](/architecture/pillars/security/sec-07-encryption/) Encryption, whose key ownership decision the classification drives
- [`SOV 3`](/architecture/pillars/sovereignty/sov-03-telemetry-residency/) Telemetry residency, which applies the classification to copies
- [`REL 8.5`](/architecture/pillars/reliability/rel-08-backup-and-restore/#rel-85-protect-backups-to-the-classification-of-their-contents) Backup protection and [`OPS 6.4`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/#ops-64-test-with-production-like-data-without-using-production-data) Test data
- [`SUS 5`](/architecture/pillars/sustainability/sus-05-data-lifecycle/) Data lifecycle, the same minimization from a different motive
