SEC 4. How do you manage identities and the lifetime of their credentials?
Last updated on
Identity is the control plane. Whoever holds a valid credential can do whatever that credential permits, which makes credential management the mechanism the rest of this pillar rests on.
The failure mode is not usually a broken authentication system. It is credentials that outlive their purpose: the key issued for a migration two years ago, the shared account nobody can retire because it is unclear what uses it, the departed employee whose access was removed from three systems out of five.
Best practices
Section titled “Best practices”SEC 4.1Federate human identity to a single source with strong authenticationSEC 4.2Give workloads their own identities rather than sharing human credentialsSEC 4.3Set an expiry on every credential, because the default is permanentSEC 4.4Prefer short-lived tokens over long-lived keys
SEC 4.1 Federate human identity to a single source with strong authentication
Section titled “SEC 4.1 Federate human identity to a single source with strong authentication”Risk if not established: High
Every system with its own user accounts is a system where joiners, movers and leavers have to be handled separately. The one that gets missed is the one that matters, and the person who left six months ago still has access to it.
A single source means one place to enforce strong authentication, one place to revoke, and one record of who exists. Revocation is the property that justifies it on its own: disabling one identity should remove access everywhere, and that is only true if everywhere federates.
Strong authentication is not optional at this point. Passwords alone fail to credential stuffing and phishing, and both are commodity attacks rather than sophisticated ones.
Enumerate the exceptions rather than assuming there are none. Local accounts on databases, admin accounts on appliances, break-glass credentials and third-party tools with their own user directories all exist in most estates, and each is an identity outside the source of truth.
On STACKIT. Where your organization has its own identity provider, federating the platform’s
identity model to it is the arrangement
that gives you the single revocation point. That decision has a sovereignty dimension as well as a
security one, since an identity provider outside the EU is a dependency at the most privileged
point of the system, which is SOV 9.
The STACKIT IdP federates with external providers over generic OIDC 1.0 , SAML 2.0 and SCIM, with a documented path for Microsoft Entra ID. Claims mapping covers user attributes such as the unique identifier, email and name. Setting up a federation goes through a support request, so it belongs in the plan early rather than in the week it is needed.
Whether group or role claims can drive role assignment is worth settling before a joiner and leaver process is designed around it: the federation guide documents claims mapping for user attributes. Plan on role assignment being a step of its own, separate from joining, and treat anything better as a saving rather than as the design.
Tradeoffs. Reliability. A single identity source is a single point of failure on the
recovery path, which REL 3.3 and REL 9.4 both identify. Break-glass access exists for that
reason and is covered in SEC 5.3.
Verify. List every place a human can authenticate in your environment. How many federate to your identity provider, and what happens in the others when someone leaves?
SEC 4.2 Give workloads their own identities rather than sharing human credentials
Section titled “SEC 4.2 Give workloads their own identities rather than sharing human credentials”Risk if not established: High
A workload using a person’s credentials inherits that person’s permissions, breaks when they leave, and makes the audit trail useless because every action appears to have been taken by them.
Workload identities solve all three. Each has its own permissions scoped to what it does, its lifecycle is independent of any person, and its actions are attributable.
Give each workload its own rather than sharing one across several. A shared identity has the union of everything its consumers need, which is more than any of them needs, and its compromise reaches all of them. It also cannot be rotated without coordinating every consumer, which is why shared credentials are the ones that never get rotated.
Pipeline identities deserve particular care. They can change production, they are frequently more
privileged than any human account, and they are used constantly. Everything in SEC 5 applies to
them, and OPS 4.1 explains why the deployment path depends on them existing.
On STACKIT. Service accounts are the workload identity mechanism, with documented authentication flows and a guide to accessing a service with one .
The product-specific roles described in roles and
permissions ,
in the form product-name.noun such as postgres-flex.admin, are mostly intended for service
accounts. That is the mechanism for giving a workload exactly the service access it
needs rather than a general Editor role, and it is the practical route to SEC 5.1.
Service account
federation
is worth reading before issuing keys, since federating an existing workload identity avoids
creating a long-lived credential at all, which is SEC 4.4.
For container images, robot accounts in Container Registry are the equivalent for automated pull and push.
Tradeoffs. Operational Excellence. More identities to create, scope and retire. Automating
their creation as part of OPS 3 is what keeps that manageable.
Verify. List the identities your workloads authenticate as. How many are personal accounts, and how many are shared between more than one workload?
SEC 4.3 Set an expiry on every credential, because the default is permanent
Section titled “SEC 4.3 Set an expiry on every credential, because the default is permanent”Risk if not established: High
A credential with no expiry is valid until someone remembers to remove it. Over a few years, that mechanism has a completion rate close to zero, which is why estates accumulate keys nobody can account for.
Expiry converts revocation from an action somebody must take into an event that happens. That is the entire argument, and it is why an expiring credential is safer than a longer one even when the longer one is better protected.
Choose the period from how often the consumer can tolerate rotation, not from how long you would like the credential to last. A ninety-day key that nobody has automated rotation for will be renewed in a hurry every ninety days, and eventually someone will set the next one to never expire to stop the interruptions. Automate the rotation first, then shorten the period.
Alert before expiry rather than at it. A credential that expires unnoticed is an outage, and it is the reason teams disable expiry after being burned once.
On STACKIT. Service account
keys
take an expiry date at creation, and it is optional: a key created without one does not
expire. The convenient path and the safe path differ here, which is what makes the default worth
setting in your own definitions under OPS 3 rather than leaving it to whoever creates the key.
Rotation is the user’s responsibility rather than an automated process, and the pattern is a short
one: create a new key, move consumers to it, delete the old one.
That is a straightforward sequence and it is one you have to build, which is a good candidate for
OPS 10.2.
Keys come in two forms. A service-generated key returns the private key once, which must then be stored securely. A user-provided key means you supply an RSA 2048 public key and the private key never leaves your control, which is the stronger option where you can manage it.
Tradeoffs. Reliability. An expiring credential is an outage waiting for the day nobody rotated it, which is why the alerting and the automation come first.
Verify. List your service account keys. How many have an expiry set, and for those that do, what happens automatically when the date approaches?
SEC 4.4 Prefer short-lived tokens over long-lived keys
Section titled “SEC 4.4 Prefer short-lived tokens over long-lived keys”Risk if not established: Medium
A long-lived key is a secret that must be stored, distributed, rotated and protected for its whole life. A short-lived token obtained from a longer-lived identity narrows the window in which a stolen credential is useful, from months to minutes.
The pattern is to hold as few long-lived secrets as possible and exchange them for short-lived tokens at the point of use. Where the workload’s environment can attest to its identity directly, the long-lived secret disappears entirely, which is the strongest version.
This is the same reasoning as expiry in SEC 4.3, applied at a different timescale. Expiry limits
how long a credential exists; short-lived tokens limit how long a leaked one works.
The remaining long-lived credentials are then a small, known set rather than an unbounded
population, which makes SEC 9 tractable.
On STACKIT. Service accounts issue short-lived access tokens, and the token how-to covers obtaining one. The key authenticates to get a token; the token authorizes the work.
For workloads on Kubernetes Engine there is a direct implementation of this best practice. Workload identity federates the cluster with the STACKIT IdP over OIDC: a managed webhook injects a token into the pod, the workload exchanges its Kubernetes service account token for a short-lived scoped STACKIT token, and no static key exists anywhere. Trust is established once, by federating the cluster with the IdP and mapping a Kubernetes service account, identified by its namespace, to a STACKIT service account.
This flow shows how Kubernetes workload identity is exchanged for a short-lived STACKIT access token:
The
stackit-pod-identity-webhookintercepts the Pod’s creation request and, based onServiceAccountannotations, injects aprojectedvolume containing a signed JWT (theServiceAccounttoken) and corresponding environment variables into the Pod spec.The application’s SDK reads this injected token volume and sends the K8s
ServiceAccounttoken to the STACKIT IdP.The STACKIT IdP validates the token signature using the cluster’s Public OIDC Discovery endpoint (JWKS) and checks the configured assertions (for example, if the cluster identity is allowed to federate the requested IdP identity).
If the validation is successful the IdP returns a temporary scoped STACKIT access token.
The application uses the returned temporary access token to call the STACKIT API.
In practice, this keeps credentials short-lived and avoids static secrets in your workloads.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
That removes the long-lived key rather than shortening its life, which is the strongest version of this best practice. One documented constraint is that the federation cannot be configured until the cluster has finished provisioning.
When using the STACKIT SDK for Go, the SDK automatically detects the injected environment variables and handles the token exchange with the STACKIT IdP behind the scenes.
For details, see the example workload provided in the stackitcloud/stackit-pod-identity-webhook GitHub repo.
Support for the STACKIT SDK for Python and Java is planned and will be available once those SDKs leave beta.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
The same distinction appears in human cluster access. Of the two kubeconfig
types ,
the dynamic one requests short-lived credentials through the CLI and contains no secret, while the
static one contains a secret and expires after at most 180 days. The first is this best practice;
the second is SEC 4.3.
Compute Engine has its own version. Attaching a service account to a server lets applications on that instance authenticate to STACKIT APIs without handling long-lived credentials, retrieving temporary tokens through the Metadata Service. Anything that reaches the Metadata Service on the instance can obtain those tokens, so the server’s own access control is what bounds it.
Elsewhere, service account federation is the equivalent for workloads that already hold an identity outside STACKIT.
Token lifetime is what decides how much the short-lived pattern actually buys, so read it off the tokens your own workload receives and build the refresh strategy on the figure you measured. A lifetime you assumed is the one that expires in the middle of a release.
Tradeoffs. Operational Excellence. Token exchange adds a step and a failure mode: a
workload that cannot reach the token endpoint cannot work, which puts that endpoint on the
availability path under REL 1.3. Performance Efficiency. A token acquired per request rather
than cached is a round trip on the hot path.
Verify. For each workload, what long-lived credential does it hold? Which of those could be replaced with federation or a short-lived token?
Related
Section titled “Related”SEC 5Least privilege, which determines what each identity may doSEC 9Secrets, which is where the remaining long-lived credentials liveSEC 11Detection and response, which watches for credential misuseSOV 9Identity sovereignty, the same federation decision from a jurisdictional angleOPS 4.1Deployment automation, whose pipeline identity is the most privileged one