SEC 5. How do you grant least privilege and keep it least over time?
Zuletzt aktualisiert am
The first half of this is well understood and the second half is where estates fail. Permissions are granted for a reason, the reason expires, and nothing removes them. Entitlement growth is silent, one-directional, and eventually produces an environment where the intended access model and the actual one have nothing in common.
The practical test is to ask, for any identity: what would this reach if its credentials were stolen tonight? If the answer is materially more than what it uses, the gap is your exposure.
Best practices
Section titled “Best practices”SEC 5.1Grant the narrowest role that permits the workSEC 5.2Prefer elevation that expires over standing privilegeSEC 5.3Make break-glass access fast, bounded and reviewedSEC 5.4Review actual entitlements against intended ones on a cadence
SEC 5.1 Grant the narrowest role that permits the work
Section titled “SEC 5.1 Grant the narrowest role that permits the work”Risk if not established: High
Broad roles are granted because they are quicker to reason about and because narrowing them takes time nobody has. The cost arrives later, when a compromised identity reaches everything rather than the part it needed.
Two dimensions to narrow, and both matter. What, which is the role: read where read suffices,
a service-specific role rather than a general one. Where, which is the scope, and SEC 2.4
explains why that one is harder to correct afterwards.
Start from what the identity actually does rather than from what it might need. The reflex to grant a broader role in case of future need is what produces the accumulation, and future need is easier to grant than past over-grant is to remove.
Human and workload identities want different treatment. Humans need enough to do varied work and
benefit most from SEC 5.2. Workloads do a fixed set of things and should be scoped tightly at
creation, since their needs do not vary.
On STACKIT. Three kinds of role exist, and the order to try them in is narrowest first.
The product-specific
roles
in the form product-name.noun, such as postgres-flex.admin, are the ones to reach for. They
are mostly intended for service accounts, which is the case where tight scoping is both most
valuable and easiest, since a workload’s needs are fixed and known.
Custom roles are the fallback where no product role fits and a general role is materially broader than the work requires. Every one of them is something to maintain, so the bar is that gap being real.
The general roles are Owner, Editor and Reader. Only Owner can assign roles, which is what lets an identity widen its own access, so the step from Editor to Owner is a different class of grant rather than a larger one. Editor covers most day-to-day work.
A product role can unlock more than its name suggests. In Kubernetes
Engine ,
owner, editor, ske.admin and ske.editor all yield an administrative kubeconfig with full
rights inside every cluster in the project, while ske.reader and reader can only download the
IdP kubeconfig and receive whatever Kubernetes RBAC grants them. The step from reader to editor is
therefore not incremental. It is the difference between namespace-scoped access and cluster
administrator, and it removes the in-cluster boundary described in SEC 2.1.
Tradeoffs. Operational Excellence. Narrow roles mean more grants and more requests, which
is friction. SEC 5.2 is how that friction stops becoming a bottleneck, and the
Security tradeoffs explain what happens when it does not.
Verify. List who holds Owner in your organization and in each project. For each, is role assignment part of what they actually do?
SEC 5.2 Prefer elevation that expires over standing privilege
Section titled “SEC 5.2 Prefer elevation that expires over standing privilege”Risk if not established: High
Most privileged access is needed occasionally and held permanently. That is the accumulation problem in its usual form: someone needed elevated access for a migration in March and still has it.
Time-bound elevation inverts the default. Day to day, an engineer holds the access their normal work needs. When something more is required, they request it, get it for a bounded period, and it expires without anyone deciding to remove it. The revocation happens because time passes rather than because a process worked.
Two properties decide whether it is adopted. It must be fast, because elevation that takes a day will be replaced by standing access that takes none. And it must be recorded, so that the pattern of elevations is visible and reviewable.
Watch what the request produces. Elevation to the same broad role every time has moved the problem rather than solved it; useful elevation is scoped to the task.
On STACKIT. Role bindings are assigned and removed through the role assignment mechanism , the Portal, the CLI and the API.
Build the time limit rather than assuming one. Design the elevation as an automated grant with an
automated revocation, which is a straightforward candidate for OPS 10.2 and gives you the expiry
whatever the binding itself supports. An elevation whose end depends on somebody remembering is
standing access with extra steps.
Every grant and removal appears in the
audit log , which is what makes the elevation
pattern reviewable under SEC 5.4 and gives an investigation the sequence of who had what and
when.
Tradeoffs. Reliability. An elevation mechanism on the recovery path is a dependency: if
elevation requires a system that is down, the recovery stalls. SEC 5.3 is the answer.
Operational Excellence. Building and operating the mechanism is real work.
Verify. How many people hold standing privileged access, and how many of them used it in the last month? What would it take for one of them to obtain the same access on demand instead?
SEC 5.3 Make break-glass access fast, bounded and reviewed
Section titled “SEC 5.3 Make break-glass access fast, bounded and reviewed”Risk if not established: Medium
Emergency access that bypasses normal controls is a security liability and a reliability necessity at the same time. Removing it means recovery stalls when the normal path is unavailable; leaving it unmanaged means a standing bypass of everything else in this pillar.
Four properties make it acceptable:
Fast. If it is slower than working around it, people will work around it, and the workaround is the one nobody logs.
Bounded. Scoped to what an emergency actually requires and expiring quickly. Break-glass that grants everything permanently is just standing privilege with a dramatic name.
Heavily logged. Every use recorded in a way the user cannot alter, which is SEC 11.1.
Reviewed after every use. Not to assign blame, which OPS 1.2 addresses, but to establish
whether the emergency was genuine and whether the normal path should have sufficed. Break-glass
used routinely is a signal that the normal path is inadequate.
Test it. An emergency path that has never been exercised has the same defect profile as any other
unrehearsed procedure, per OPS 8.3, and it will be discovered during the incident.
On STACKIT. The mechanism is a deliberately separated identity with a defined elevation path, constructed from roles and permissions rather than a dedicated platform feature.
Its use is recorded in the audit log like any
other action, which gives the after-the-fact review its evidence. Note the 90-day retention in the
Portal: if break-glass review happens quarterly, that window is tighter than it looks, and
exporting through Telemetry Router under OPS 7.4 is what extends it.
The recovery-path point from REL 9.4 applies directly. If break-glass depends on the identity
provider being reachable, and the identity provider is inside the failure domain, the mechanism
does not exist when it is needed.
Tradeoffs. Security, directly. Break-glass is a deliberate weakening of the access model, justified by the reliability requirement. The bounding and the logging are what keep the trade acceptable rather than making it free.
Verify. What is your break-glass procedure, when was it last used, and what did the review after that use conclude? When was it last tested?
SEC 5.4 Review actual entitlements against intended ones on a cadence
Section titled “SEC 5.4 Review actual entitlements against intended ones on a cadence”Risk if not established: High
The intended access model and the actual one diverge continuously, and only a review that looks at the actual one will notice.
Review the effective permissions rather than the grants. In an inheriting hierarchy, what an
identity can reach is the union of every binding on every ancestor, so a project-level review that
ignores organization and folder bindings will report an access model that does not exist. That is
the direct consequence of SEC 2.4.
Three categories to look for, in order of how much they usually find:
- Identities nobody recognizes. Service accounts from projects that ended, keys issued for migrations, accounts for people who left.
- Permissions nobody uses. Granted for a reason that has passed. Actual usage is the evidence here rather than the grant.
- Grants at the wrong scope. Applied at organization or folder level because it was convenient at the time, now reaching projects that did not exist then.
Have the review done by someone who can say whether the access is still needed, which is usually the owning team rather than a central function. A review nobody with context performs will approve everything.
On STACKIT. Effective access is assembled from bindings at organization, folder and project scope, and the roles and permissions model makes clear that ancestors contribute. The review therefore starts at the top of the hierarchy rather than at the project.
The audit log records actions taken, which is
the usage evidence that distinguishes a permission in use from one merely held. Its Portal
retention, which OPS 7.4 covers, bounds how far back that evidence reaches unless it has been
exported.
Bindings are queryable through the API
and CLI , which is what makes the review
a scheduled report rather than a manual inspection, and a good candidate for OPS 10.2.
Tradeoffs. Operational Excellence. Reviews cost time from the people best placed to judge, which is also the people with the least time. Automating the report so only the judgement is manual is what makes the cadence survivable.
Verify. When did you last review who can access production, what did it change, and did it include bindings inherited from folders and the organization?
Related
Section titled “Related”SEC 2.4Scope, which is why breadth is expensive to correctSEC 4Identity, which supplies the principals these grants attach toSEC 11Detection and response, which watches for use of what was grantedSOV 7Auditability, which needs the same records for a different purposeOPS 3.2Reviewing infrastructure changes, where role bindings should also be defined