SEC 8. How do you harden resources and keep them patched?
Zuletzt aktualisiert am
Attackers overwhelmingly use known vulnerabilities in unpatched software, because it works and requires no research. That makes patching the highest-return security activity available and the least interesting to do.
Hardening is the other half: removing what does not need to be there, so that there is less to patch and less to exploit. A component reduced to what it needs has a smaller attack surface permanently, without anyone having to react to a disclosure.
Best practices
Section titled “Best practices”SEC 8.1Reduce each component to what it needs to runSEC 8.2Define a patch cadence with a maximum exposure window per severitySEC 8.3Patch the whole stack, not only the operating systemSEC 8.4Rebuild rather than patch in place where you can
SEC 8.1 Reduce each component to what it needs to run
Section titled “SEC 8.1 Reduce each component to what it needs to run”Risk if not established: Medium
Default configurations are built for compatibility rather than for security, which means they enable more than any specific workload needs. Every enabled feature is attack surface and something that must be patched.
The reductions that pay:
- Remove packages and services that are not used. A container image with a shell and a package manager gives an attacker tools they would otherwise have to bring.
- Close ports that nothing listens on legitimately, which
SEC 6.1handles at the network layer and hardening handles at the host. - Drop privileges. Processes running as root, containers running privileged, and mounts that do not need to be writable are all avoidable defaults.
- Disable unused authentication methods, particularly ones that predate your current standard.
Do it in the image or the definition rather than after deployment. A hardening script that runs
post-deployment can fail, be skipped, or be absent from an instance created outside the normal
path. OPS 3.1 is what makes hardening reproducible.
On STACKIT. Security hardening provides platform-specific guidance, split into Compute Engine and networks . Starting from that rather than from a generic benchmark saves the part of the work that is platform-dependent.
CSPM measures the result against benchmarks
including BSI C5 Basic and the STACKIT Standard, which turns hardening from a one-time exercise
into a checked state under SEC 1.2.
For Kubernetes, enhancing cluster security covers the cluster-level configuration.
Tradeoffs. Operational Excellence. A minimal image is harder to debug, since the tools you would reach for are absent. That is the point and it is genuinely inconvenient; ephemeral debug containers are the usual answer for Kubernetes.
Verify. Take one production image or instance. What is installed that the workload does not use? What privileges does the main process hold that it does not need?
SEC 8.2 Define a patch cadence with a maximum exposure window per severity
Section titled “SEC 8.2 Define a patch cadence with a maximum exposure window per severity”Risk if not established: High
Patching without a stated window happens when someone has time, which under delivery pressure is never. A defined maximum exposure period per severity turns it into a commitment that can be measured and missed visibly.
Set the windows from severity and exposure together. A critical vulnerability in an internet-reachable component is a different urgency from the same vulnerability in an internal batch job, and treating them identically means either the first is too slow or the second is disruptive for no reason.
Measure from disclosure rather than from when you noticed. The exposure window is the period an exploit exists and you are vulnerable, and your discovery date is not part of the attacker’s timeline. That also makes the detection latency visible, which is usually where the time goes.
Have a path for the emergency case. Some vulnerabilities warrant patching outside the normal cycle, and that decision needs an owner before it is needed rather than during it.
Track exceptions. A component that cannot be patched needs a compensating control and a recorded
acceptance under REL 3.4, not silence.
On STACKIT. Server Update
Management provides
automated operating system updates for Linux and Windows with configurable
schedules .
That covers the largest recurring patching burden on Compute Engine and removes it from OPS 10
at the same time.
Scheduling pre-production ahead of production, as OPS 6.3 suggests, turns patching into a change
that passes through your deployment stages rather than a separate class of work.
For managed services, patching of the underlying platform is STACKIT’s responsibility, which is the usual division and one of the reasons managed services reduce this burden. What remains yours is the version you run: SKE version updates documents the cluster version lifecycle, and staying on a supported version is a customer-side decision with a security consequence.
If the Kubernetes or OS version of a cluster has reached its expiration date, SKE starts a mandatory update to the highest available patch version of the current minor version, or to the highest patch version of the consecutive minor version that is not classified as preview version. Note that mandatory version updates run, even if the auto update for the Kubernetes version is deactivated since using a supported version is crucial for your cluster’s security and stability. To plan ahead, you can check the exact expiration dates for your current Kubernetes and node pool versions in the cluster.status.expiration field via the SKE API.
Was ist das?
Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.
Each managed service publishes a lifecycle reference with dated end-of-support per version, such as PostgreSQL Flex .
Diesen Abschnitt gibt es nur auf Englisch.
- 1 major release per year
- Current version is supported for at least 180 days after the release of the next major version
Was ist das?
Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.
Diesen Abschnitt gibt es nur auf Englisch.
| Name | Major version | Release date | End of life date |
|---|---|---|---|
| STACKIT PostgreSQL Flex | 18 (Beta) | September 2026 | - |
| STACKIT PostgreSQL Flex | 17 | September 2026 | November 2029 |
| STACKIT PostgreSQL Flex | 16 | September 2024 | November 2028 |
| STACKIT PostgreSQL Flex | 15 | May 2023 | November 2027 |
| STACKIT PostgreSQL Flex | 14 | May 2023 | November 2026 |
| STACKIT PostgreSQL Flex | 13 | May 2023 | November 2025 |
Was ist das?
Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.
Those dates are the input to a version upgrade plan, and they are known far enough ahead that reaching end of support is a scheduling failure rather than a surprise.
Tradeoffs. Reliability. Patching is change, and change causes incidents. The resolution is
the same as OPS 4.3: small, frequent, staged patching is safer than large infrequent patching,
and both are safer than not patching.
Verify. What is your maximum exposure window for a critical vulnerability, measured from disclosure? For the last critical one, what was the actual figure?
SEC 8.3 Patch the whole stack, not only the operating system
Section titled “SEC 8.3 Patch the whole stack, not only the operating system”Risk if not established: High
Operating system patching is automated in most estates. The layers above it frequently are not, and that is where a growing share of exploited vulnerabilities live.
The layers that need their own answer:
- Application dependencies. Libraries pulled in at build time, including the transitive ones nobody chose.
- Container base images. A base image is an operating system that your OS patching does not touch, because it is inside the artefact.
- Runtimes and frameworks. Language runtimes, web servers, and the frameworks between them.
- Infrastructure components you operate. An ingress controller, a mesh, a monitoring agent.
- The pipeline itself. Build tooling and its plugins, which have production access.
Each needs a mechanism, and the mechanism is usually rebuild rather than update in place, which is
SEC 8.4.
On STACKIT.
Container Registry
provides vulnerability scanning for stored images, which covers the base image and dependency
layers of anything you push. That makes the finding automatic; acting on it is the part you build,
and it belongs in the pipeline under SEC 10.2.
Dependency scanning at source level runs in STACKIT Pipelines using whichever scanner you choose, and its GitHub Actions compatibility means existing scanning actions can be reused rather than reimplemented.
Tradeoffs. Operational Excellence. Continuous dependency updating is a steady stream of changes, each needing testing. Automating the update and letting the test suite gate it is the only version that scales, and it depends on the test suite being trustworthy.
Verify. For one production service, when were its dependencies, its base image and its runtime last updated? Which of those has an automated mechanism?
SEC 8.4 Rebuild rather than patch in place where you can
Section titled “SEC 8.4 Rebuild rather than patch in place where you can”Risk if not established: Medium
Patching in place accumulates state. Instances that have run for a year carry the residue of every change ever applied, they drift from each other, and no two are quite the same, which makes them individually unpredictable.
Rebuilding from a current definition removes all of that. The new instance has exactly what the definition says, it matches its peers, and the patching problem becomes a build problem, which is already automated.
This requires the workload to tolerate instances being replaced, which is the same property REL 4 needs for zone distribution and OPS 4 needs for deployment. Where it holds, patching is a
deployment rather than an operation.
Where it does not hold, in-place patching remains legitimate, and those components deserve the attention: they are the long-lived ones that drift, and they are usually the ones holding state.
Note the interaction with OPS 3.3: rebuilt infrastructure cannot drift, because it is recreated
from the definition each time. That is a security property as much as an operational one.
On STACKIT. Image-based rebuild for Compute Engine is orchestrated through the
API ,
CLI or Terraform
provider ,
which is the same path OPS 3 uses.
For Kubernetes, node replacement during version
updates is
the mechanism, and it exercises the same properties REL 10.4 wants rehearsed. A cluster upgrade
is a scheduled resilience test that happens anyway.
Where rebuild is not possible, Server Update Management is the in-place path and remains the right answer for stateful long-lived servers.
Tradeoffs. Reliability. Rebuilding is more disruptive per patch than updating in place,
and requires the redundancy from REL 4.1 to be non-disruptive overall. Cost Optimization.
Brief overlap of old and new capacity during replacement.
Verify. How long has your longest-running production instance been running, and how many times has it been patched in place? Could it be replaced rather than patched?
Related
Section titled “Related”SEC 1Security baseline, which states the hardened configurationSEC 10Supply chain, which shares the dependency and image scanningOPS 6.3Version alignment, which patching should followOPS 10.2Toil elimination, which Server Update Management addresses directlyREL 3.1Failure modes, which planned maintenance belongs in