---
id: SEC09
pillar: security
title: SEC 9. How do you store and rotate application secrets?
description: A secret in a repository is a secret in every clone, every backup and every fork of that repository, permanently. Where they belong and how they get replaced.
status: draft
services: [secrets-manager, kms]
sidebar:
  order: 18
  label: Secrets
source_url: "https://framework.stackit.cloud/architecture/pillars/security/sec-09-secrets/"
source_file: "docs/architecture/pillars/security/sec-09-secrets.mdx"
---

Secrets are the credentials your workload needs in order to work: database passwords, API keys,
signing keys, certificates. They are attractive to an attacker precisely because they are the
shortest path from a foothold to the data.

The characteristic failure is not a broken secret store. It is a secret in a place nobody thought
of as a store: a repository, an image layer, an environment file, a CI variable, a wiki page, or a
chat message from 2022.

## Best practices

- [`SEC 9.1`](/architecture/pillars/security/sec-09-secrets/#sec-91-keep-secrets-in-a-purpose-built-store-never-in-code-or-images) Keep secrets in a purpose-built store, never in code or images
- [`SEC 9.2`](/architecture/pillars/security/sec-09-secrets/#sec-92-give-each-consumer-its-own-secret-rather-than-sharing-one) Give each consumer its own secret rather than sharing one
- [`SEC 9.3`](/architecture/pillars/security/sec-09-secrets/#sec-93-rotate-on-a-schedule-and-after-any-suspected-exposure) Rotate on a schedule and after any suspected exposure
- [`SEC 9.4`](/architecture/pillars/security/sec-09-secrets/#sec-94-detect-secrets-that-have-leaked-into-places-they-should-not-be) Detect secrets that have leaked into places they should not be

---

## SEC 9.1 Keep secrets in a purpose-built store, never in code or images

**Risk if not established:** High

A secret committed to a repository is present in every clone, every fork, every backup and the
history, permanently. Removing it from the current version does not remove it, which is why the
only correct response to a committed secret is to rotate it.

Container images have the same property with less visibility. A secret in a build layer survives
even if a later layer deletes it, and anyone who can pull the image can extract it.

The places that count as unsafe, and are commonly used anyway: source repositories, container
image layers, environment files committed alongside code, CI variables that are not backed by a
secret store, configuration management repositories, and infrastructure definitions.

The last one deserves a note, because [`OPS 3.1`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-31-define-every-production-resource-in-version-control) asks for everything to be in code. The resolution
is that the *existence* of a secret and its consumers belong in code, while the *value* lives in
the store and is referenced. That distinction is what lets both practices hold at once.

A purpose-built store gives you what a file cannot: access control per secret, an audit trail of
reads, versioning, and a rotation path.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/security/secrets-manager/">Secrets Manager</LinkChip>
is the managed store, with
<LinkChip href="https://docs.stackit.cloud/products/security/secrets-manager/getting-started/versioning-in-secrets-manager/">versioning</LinkChip>
that makes rotation possible without a simultaneous cutover, and <LinkChip href="https://docs.stackit.cloud/products/security/secrets-manager/getting-started/configure-the-secrets-manager/">configuration
guidance</LinkChip>
for setting it up.

Distinguish it from <LinkChip href="https://docs.stackit.cloud/products/security/kms/">KMS</LinkChip>, which manages
cryptographic keys rather than application secrets. KMS keys perform operations without the key
material leaving the service; Secrets Manager stores values your application retrieves. Different
purposes, and using one for the other is a common source of confusion.

Retrieval at runtime rather than at build time is the property that matters, since it keeps the
value out of the artefact.

**Tradeoffs.** **Reliability.** The secret store is now on the startup path of every workload that
reads from it, which puts it in the availability composition under [`REL 1.3`](/architecture/pillars/reliability/rel-01-reliability-targets/#rel-13-check-every-target-against-the-published-availability-of-the-services-it-depends-on). Caching retrieved
secrets in memory is the usual mitigation and it lengthens the window in which a rotated secret is
still in use.

**Verify.** Search your repositories and image layers for credentials. What did you find, and has
each of those been rotated since?

---

## SEC 9.2 Give each consumer its own secret rather than sharing one

**Risk if not established:** Medium

A shared secret cannot be rotated without coordinating every consumer, which is why shared secrets
are the ones that never get rotated. It also cannot be attributed: when it is misused, the audit
trail shows the secret rather than which consumer used it.

Per-consumer secrets fix both. Rotation is independent, revocation is surgical, and misuse points
somewhere specific.

The same reasoning applies to environments. A database credential shared between production and
staging means a staging compromise reaches production, which is exactly the boundary [`SEC 2.3`](/architecture/pillars/security/sec-02-segmentation/#sec-23-separate-environments-and-be-explicit-about-what-the-separation-rests-on)
built.

This is the secret-level version of [`SEC 4.2`](/architecture/pillars/security/sec-04-identity/#sec-42-give-workloads-their-own-identities-rather-than-sharing-human-credentials), and where the underlying system supports distinct
identities, that is the better mechanism: a service account per workload is stronger than a shared
password issued twice.

**On STACKIT.** Secrets Manager provides per-service isolation, and the <LinkChip href="https://docs.stackit.cloud/products/security/secrets-manager/getting-started/create-a-service-in-secrets-manager/">service and access
control
model</LinkChip>
is where the boundary is drawn.

For the platform itself, <LinkChip href="https://docs.stackit.cloud/platform/access-and-identity/service-accounts/">service
accounts</LinkChip> are the
per-consumer identity, and for image registries <LinkChip href="https://docs.stackit.cloud/products/developer-platform/container-registry/how-tos/automate-workflows-with-robot-accounts/">robot
accounts</LinkChip>
serve the same purpose for automated pull and push. Both are preferable to sharing a credential,
and both are [`SEC 4.2`](/architecture/pillars/security/sec-04-identity/#sec-42-give-workloads-their-own-identities-rather-than-sharing-human-credentials) applied.

**Tradeoffs.** **Operational Excellence.** More secrets to create, distribute and retire. That is
manageable when their creation is part of the automated provisioning in [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) and painful when
it is manual.

**Verify.** List your secrets and their consumers. How many are used by more than one workload or
more than one environment?

---

## SEC 9.3 Rotate on a schedule and after any suspected exposure

**Risk if not established:** High

An unrotated secret has an unbounded exposure window. Every person who ever had access to it still
has it, every place it was ever copied still holds it, and a compromise from two years ago is
still live.

Rotation bounds that. It also, and this matters more in practice, proves that rotation works. A
secret that has never been rotated has an untested rotation path, which will be discovered during
the incident where rotation is urgent.

Two triggers. A **schedule**, set from the consequence of exposure rather than from a round
number, and **any suspected exposure**: a departure, a commit, a compromise, a third party
incident, or uncertainty about whether a secret leaked. Uncertainty is sufficient reason, because
the cost of rotating unnecessarily is low and the cost of not rotating a leaked secret is not.

Design for overlap. The version that works is: create the new secret, let both be valid, move
consumers, then retire the old one. Rotation that requires a simultaneous cutover is rotation that
causes an outage, which is why it gets deferred.

Automate it. Manual rotation happens until it is inconvenient, then stops.

**On STACKIT.** <LinkChip href="https://docs.stackit.cloud/products/security/secrets-manager/getting-started/versioning-in-secrets-manager/">Secrets Manager
versioning</LinkChip>
is what supports the overlap pattern, since a new version can exist while consumers still read the
previous one.

For platform credentials, <LinkChip href="https://docs.stackit.cloud/platform/access-and-identity/service-accounts/how-tos/manage-service-account-keys/">service account
keys</LinkChip>
document the rotation sequence explicitly: create a new key, update the systems using it, delete
the old one. The documentation is also explicit that rotation is the user's responsibility rather
than an automated process, so the schedule and the automation are yours to build. That is a
well-shaped candidate for [`OPS 10.2`](/architecture/pillars/operational-excellence/ops-10-toil-elimination/#ops-102-automate-what-recurs-starting-with-the-riskiest-rather-than-the-most-frequent).

The optional expiry described in [`SEC 4.3`](/architecture/pillars/security/sec-04-identity/#sec-43-set-an-expiry-on-every-credential-because-the-default-is-permanent) is the forcing function: a key with an expiry date
makes rotation a deadline rather than an intention.

Some services document their own rotation path, such as <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/rotate-ske-credentials/">rotating SKE
credentials</LinkChip>.
Where a workload uses <LinkChip href="https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/workload-identity/">workload
identity</LinkChip>
instead, there is no static credential to rotate at all, which is the outcome [`SEC 4.4`](/architecture/pillars/security/sec-04-identity/#sec-44-prefer-short-lived-tokens-over-long-lived-keys) aims at.

**Tradeoffs.** **Reliability.** Rotation is change, and a botched rotation is an outage. The
overlap pattern is what makes it safe, and it requires the consuming systems to tolerate two valid
credentials, which is a design property rather than a configuration.

**Verify.** For your most privileged secret, when was it last rotated and by what mechanism? If
you had to rotate it in the next hour, what would break?

---

## SEC 9.4 Detect secrets that have leaked into places they should not be

**Risk if not established:** Medium

Secrets end up in the wrong places despite the rules, because the rules are followed by people
under time pressure. Detection is what turns that from a permanent exposure into a bounded one.

Scan in three places, because they catch different mistakes. **Pre-commit**, on the developer's
machine, which prevents the commit entirely and is the cheapest. **In the pipeline**, which
catches what bypassed the first. And **across existing repositories and their history**, which
finds what was committed before any of this was in place, and which is usually where the first
real findings are.

Include the places that are not repositories: image layers, logs, configuration stores, and issue
trackers. Logs are a recurring source, since a debug statement that dumps a request or a
configuration object will happily include a credential.

Treat every finding as a real exposure. A secret found in a repository has been in every clone of
that repository, and assuming it was not read is not a defensible position. Rotate first, then
investigate.

**On STACKIT.** Scanning runs in your pipeline rather than being a platform feature. <LinkChip href="https://docs.stackit.cloud/products/developer-platform/git/basics/stackit-pipelines/">STACKIT
Pipelines</LinkChip>
is GitHub Actions compatible, so existing secret-scanning actions can be reused rather than
reimplemented, which makes this one of the cheaper controls to add.

<LinkChip href="https://docs.stackit.cloud/products/developer-platform/container-registry/">Container Registry</LinkChip>
scans images for vulnerabilities, which is a different check from secret detection. Both belong in
the artefact path and neither substitutes for the other.

For logs, the control is preventive rather than detective: structured logging with explicit field
selection under [`OPS 7.1`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-71-emit-metrics-logs-and-traces-and-make-them-joinable), so that a credential cannot arrive in a log line by accident.

**Tradeoffs.** **Operational Excellence.** Secret scanners produce false positives, and a scanner
that blocks on noise gets bypassed. Tuning is required, and it is worth the effort because the
true positives are severe.

**Verify.** Is secret scanning running on your repositories, including their history? What did the
last full scan find, and was each finding rotated?

---

## Related

- [`SEC 4.3`](/architecture/pillars/security/sec-04-identity/#sec-43-set-an-expiry-on-every-credential-because-the-default-is-permanent) Credential expiry, the same discipline for platform credentials
- [`SEC 7`](/architecture/pillars/security/sec-07-encryption/) Encryption, and specifically the distinction between KMS and a secret store
- [`SEC 10`](/architecture/pillars/security/sec-10-supply-chain/) Supply chain, which shares the pipeline scanning
- [`OPS 3.1`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/#ops-31-define-every-production-resource-in-version-control) Everything as code, and where secrets are the deliberate exception
- [`SEC 11`](/architecture/pillars/security/sec-11-detection-and-response/) Detection and response, which acts on a confirmed leak
