---
id: SEC11
pillar: security
title: SEC 11. How do you detect security events and respond to them?
description: Under assume breach the question is not whether someone gets in but how long before you notice. Signals an attacker cannot edit, and a rehearsed response.
status: draft
services: [observability, cspm]
sidebar:
  order: 20
  label: Detection & response
source_url: "https://framework.stackit.cloud/architecture/pillars/security/sec-11-detection-and-response/"
source_file: "docs/architecture/pillars/security/sec-11-detection-and-response.mdx"
---

Assume breach makes detection as important as prevention. If an intrusion is going to happen
eventually, the metric that matters is the time between it happening and someone noticing, because
that is the window in which an attacker moves laterally, escalates and exfiltrates.

The second half is the response. An incident process first executed during an incident is a
document, and security incidents add pressures that operational ones do not: evidence
preservation, legal notification duties, and the possibility that your own infrastructure cannot
be trusted.

## Best practices

- [`SEC 11.1`](/architecture/pillars/security/sec-11-detection-and-response/#sec-111-collect-security-relevant-signals-where-the-subject-cannot-alter-them) Collect security-relevant signals where the subject cannot alter them
- [`SEC 11.2`](/architecture/pillars/security/sec-11-detection-and-response/#sec-112-alert-on-patterns-that-indicate-compromise-rather-than-on-volume) Alert on patterns that indicate compromise rather than on volume
- [`SEC 11.3`](/architecture/pillars/security/sec-11-detection-and-response/#sec-113-define-the-response-before-an-alert-fires) Define the response before an alert fires
- [`SEC 11.4`](/architecture/pillars/security/sec-11-detection-and-response/#sec-114-rehearse-it-including-the-parts-that-are-not-technical) Rehearse it, including the parts that are not technical

---

## SEC 11.1 Collect security-relevant signals where the subject cannot alter them

**Risk if not established:** High

An attacker with access to a system will delete the evidence of how they got there. That is
standard practice, not sophistication, which means logs that stay on the system that produced them
are not evidence.

Ship them somewhere with a different trust boundary: a different project, different credentials,
and write access that does not include delete. The identity that produces a log should not be able
to remove it.

The signals worth collecting for security purposes are narrower than everything:

- **Authentication**, successful and failed, for humans and workloads.
- **Authorization changes.** Role grants, particularly at organization and folder scope, given the
  inheritance described in [`SEC 2.4`](/architecture/pillars/security/sec-02-segmentation/#sec-24-grant-at-the-narrowest-scope-because-breadth-cannot-be-narrowed-further-down).
- **Administrative actions** on resources: created, deleted, reconfigured.
- **Data access** to classified data sets, where the classification warrants it.
- **Network anomalies**, particularly outbound to destinations that are not on the allowlist from
  [`SEC 6.3`](/architecture/pillars/security/sec-06-network-controls/#sec-63-control-egress-because-that-is-how-data-leaves).

Retention for security purposes is longer than for operations, because the median time to detect
an intrusion is measured in months rather than days. Logs that expire before you look are logs you
did not have.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> records every
action by users, service accounts and the platform, at organization, folder and project scope, and
is enabled by default rather than requiring configuration. That covers the authentication,
authorization and administrative categories above without you building anything, which is a
substantial head start.

The property to plan around is retention: **90 days in the Portal**. For security purposes that is
short relative to typical dwell time, and the mechanism for extending it is exporting through
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/">Telemetry Router</LinkChip>
to a destination you choose, with a <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/getting-started/create-your-first-instance-connect-it-with-logs-and-query-your-audit-data/">worked
example</LinkChip>
of connecting it to Logs.

That export is the step that turns a 90-day operational record into a security record, and it has
to exist before the intrusion you will investigate. Sending it to a project with separate
credentials is what gives it the different trust boundary this best practice asks for.

Application-level security signals go to
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip> or
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/logs/">Logs</LinkChip> and depend on your
instrumentation under [`OPS 7`](/architecture/pillars/operational-excellence/ops-07-observability/). Their retention has a ceiling: the <LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/">Observability
service
plans</LinkChip>
default logs and traces to 5 days with a maximum of 30. That is a sensible operational figure and
well short of typical intrusion dwell time, which means security retention is an export decision
rather than a settings decision, in the same way the audit log is.

**Tradeoffs.** **Cost Optimization.** Long retention of security logs accumulates continuously and
its value is invisible until an investigation, which is why it gets cut. [`OPS 7.4`](/architecture/pillars/operational-excellence/ops-07-observability/#ops-74-set-retention-by-the-value-of-each-signal-rather-than-uniformly) is the argument
for recording the reason next to the number.

**Verify.** If an attacker gained administrative access to your environment today, which records
of that could they delete? How far back does your audit history actually reach?

---

## SEC 11.2 Alert on patterns that indicate compromise rather than on volume

**Risk if not established:** High

Collecting signals is necessary and produces nothing on its own. What turns collection into
detection is a small set of alerts on patterns that genuinely indicate something wrong.

The patterns that earn an alert are the ones that are rare in normal operation and characteristic
of an intrusion:

- **Authentication anomalies.** A service account authenticating from somewhere new, a burst of
  failures followed by a success, use of a credential that has been dormant.
- **Privilege escalation.** A role granted at organization scope, an identity granted Owner, a
  break-glass account used.
- **Unusual data access.** Bulk reads of a classified data set, access outside working patterns.
- **Outbound anomalies.** Traffic to destinations outside the allowlist, or volumes that do not
  fit the workload.

Keep the set small deliberately. The failure mode here is a hundred rules producing a hundred
alerts a day, at which point nobody reads any of them and the detection capability is nominal. Two
alerts a week that always mean something beat that comfortably.

Tune with the environment rather than against a template. What is anomalous depends on what is
normal, and normal is specific to your workload.

**On STACKIT.**
<LinkChip href="https://docs.stackit.cloud/products/security/cspm/">CSPM</LinkChip> detects configuration weaknesses, which
is posture rather than activity. It answers "is something exposed" and not "is someone using it".
Both are needed and they are different questions.

For activity, alerting is built on the audit records exported under [`SEC 11.1`](/architecture/pillars/security/sec-11-detection-and-response/#sec-111-collect-security-relevant-signals-where-the-subject-cannot-alter-them) and on application
signals in
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/">Observability</LinkChip>, using
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/manage-alert-groups-and-alerts/">alert
groups</LinkChip>
to separate security alerts from operational ones so that they route differently.

There is no managed security information and event management service, so
correlation rules across signals are built with the tooling you choose. That is the usual
division, and for many workloads the small set of high-value alerts described above is achievable
without a dedicated platform.

**Tradeoffs.** **Operational Excellence.** Tuning is continuous, and an untuned rule set becomes
noise within weeks. Reviewing which alerts fired and which led to action, as [`REL 10.2`](/architecture/pillars/reliability/rel-10-health-model-and-testing/#rel-102-alert-on-user-visible-symptoms-rather-than-on-component-metrics-alone) suggests,
is what keeps it honest.

**Verify.** Which security alerts exist, how many fired last month, and how many led to an
investigation? If a service account key were used from an unexpected source tonight, what would
happen?

---

## SEC 11.3 Define the response before an alert fires

**Risk if not established:** High

Security incidents share a structure with operational ones, which [`OPS 9.1`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-91-define-severity-roles-and-escalation-before-you-need-them) covers, and add
requirements that change the response.

**Evidence preservation comes before mitigation** more often than in an operational incident. The
instinct to terminate a compromised instance destroys the memory state, the process list and the
attacker's tooling. Snapshot first where the severity warrants it.

**Containment is a distinct phase.** Before eradicating, stop the spread: revoke the credential,
isolate the network segment, disable the account. Doing this well depends on the segmentation from
[`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) actually existing.

**The infrastructure may not be trustworthy.** If an attacker has administrative access, your
monitoring, your deployment pipeline and your communication channels may all be observed. A
response plan that assumes they are clean is a plan the attacker is reading.

**Legal and regulatory duties have clocks.** Personal data breaches carry notification deadlines
that begin at awareness rather than at resolution, which means legal counsel is part of the
response rather than something that follows it. [`SOV 8`](/architecture/pillars/sovereignty/sov-08-compliance-mapping/) establishes which obligations apply.

Name who decides. Isolating a production system to contain an intrusion is a costly decision, and
it needs an owner who can make it quickly.

**On STACKIT.** The <LinkChip href="https://docs.stackit.cloud/platform/audit-log/">audit log</LinkChip> is the primary
investigation source for what an attacker did on the platform, subject to the retention and export
considerations in [`SEC 11.1`](/architecture/pillars/security/sec-11-detection-and-response/#sec-111-collect-security-relevant-signals-where-the-subject-cannot-alter-them).

Containment mechanisms are the ones already covered: revoking a service account key or a role
binding under <LinkChip href="https://docs.stackit.cloud/platform/access-and-identity/">access and identity</LinkChip>, and
network isolation through <LinkChip href="https://docs.stackit.cloud/products/network/core-networking/security-groups/">security
groups</LinkChip> or the
<LinkChip href="https://docs.stackit.cloud/products/network/network-security/unified-firewall/">Unified
Firewall</LinkChip>. Knowing
which of those you would use, and having the permission to use it, belongs in the plan rather than
in the incident.

<LinkChip href="https://status.stackit.cloud">status.stackit.cloud</LinkChip> is where a platform-side event would appear,
which is worth checking early to establish whether what you are seeing is yours.

**Tradeoffs.** **Reliability.** Containment actions cause outages by design. Isolating a system
stops the spread and stops the service, and that trade is a decision somebody has to be authorized
to make.

**Verify.** Who can revoke a compromised credential in your environment right now, and how long
would it take? What is your notification deadline for a personal data breach, and who starts that
clock?

---

## SEC 11.4 Rehearse it, including the parts that are not technical

**Risk if not established:** Medium

Security incident response is rehearsed less than any other procedure in this framework, because
it is unpleasant and because the scenarios feel unlikely until they are not.

Rehearse at the depth the risk justifies. A tabletop exercise costs a few hours and finds most of
the process gaps: who decides, who is called, what is said, and which permission nobody has. A
technical exercise verifies that containment actually works and that the evidence you assumed
exists does.

Rehearse the non-technical parts, since they are where the delays are. Notification duties,
customer communication, and the question of who talks to a regulator are all decisions people make
badly under pressure and well in advance.

Include the assumption that the environment is compromised. A rehearsal where the monitoring, the
pipeline and the chat channel are all trusted is a rehearsal of the easy case.

Record what it finds and track it like any other incident action under [`OPS 9.4`](/architecture/pillars/operational-excellence/ops-09-incident-management/#ops-94-track-review-actions-to-completion). A rehearsal
whose findings are not closed produced a document rather than a capability.

**On STACKIT.** Rehearsal needs an environment, which [`OPS 6`](/architecture/pillars/operational-excellence/ops-06-environment-consistency/) and [`OPS 3`](/architecture/pillars/operational-excellence/ops-03-everything-as-code/) provide: creating a
representative environment from definitions is what makes a technical exercise affordable without
touching production.

Verify the evidence path during the exercise rather than assuming it. Whether audit records were
actually exported through
<LinkChip href="https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/">Telemetry Router</LinkChip>,
and whether anyone can query them, is the kind of thing that is configured once and never checked
until it matters.

**Tradeoffs.** **Cost Optimization.** Rehearsal time from people whose time is expensive, for a
scenario that may not occur. It is the same argument as [`REL 9.3`](/architecture/pillars/reliability/rel-09-disaster-recovery/#rel-93-rehearse-it-and-time-the-rehearsal-against-the-rto) and it has the same answer: the
alternative is a response capability nobody has evidence for.

**Verify.** When did you last rehearse a security incident, what scenario, and what did it find?
Did the rehearsal include the notification and communication steps?

---

## Related

- [`OPS 9`](/architecture/pillars/operational-excellence/ops-09-incident-management/) Incident management, whose process this extends
- [`SEC 1`](/architecture/pillars/security/sec-01-security-baseline/) Security baseline, which detects configuration weakness rather than activity
- [`SEC 2`](/architecture/pillars/security/sec-02-segmentation/) Segmentation, which containment depends on
- [`SEC 5.4`](/architecture/pillars/security/sec-05-least-privilege/#sec-54-review-actual-entitlements-against-intended-ones-on-a-cadence) Entitlement review, which uses the same records
- [`SOV 7`](/architecture/pillars/sovereignty/sov-07-auditability/) Auditability and [`SOV 8`](/architecture/pillars/sovereignty/sov-08-compliance-mapping/) Compliance mapping, which set the retention and the duties
