Skip to content
Beta

SEC 11. How do you detect security events and respond to them?

Last updated on

Assume breach makes detection as important as prevention. If an intrusion is going to happen eventually, the metric that matters is the time between it happening and someone noticing, because that is the window in which an attacker moves laterally, escalates and exfiltrates.

The second half is the response. An incident process first executed during an incident is a document, and security incidents add pressures that operational ones do not: evidence preservation, legal notification duties, and the possibility that your own infrastructure cannot be trusted.

  • SEC 11.1 Collect security-relevant signals where the subject cannot alter them
  • SEC 11.2 Alert on patterns that indicate compromise rather than on volume
  • SEC 11.3 Define the response before an alert fires
  • SEC 11.4 Rehearse it, including the parts that are not technical

SEC 11.1 Collect security-relevant signals where the subject cannot alter them

Section titled “SEC 11.1 Collect security-relevant signals where the subject cannot alter them”

Risk if not established: High

An attacker with access to a system will delete the evidence of how they got there. That is standard practice, not sophistication, which means logs that stay on the system that produced them are not evidence.

Ship them somewhere with a different trust boundary: a different project, different credentials, and write access that does not include delete. The identity that produces a log should not be able to remove it.

The signals worth collecting for security purposes are narrower than everything:

  • Authentication, successful and failed, for humans and workloads.
  • Authorization changes. Role grants, particularly at organization and folder scope, given the inheritance described in SEC 2.4.
  • Administrative actions on resources: created, deleted, reconfigured.
  • Data access to classified data sets, where the classification warrants it.
  • Network anomalies, particularly outbound to destinations that are not on the allowlist from SEC 6.3.

Retention for security purposes is longer than for operations, because the median time to detect an intrusion is measured in months rather than days. Logs that expire before you look are logs you did not have.

On STACKIT. The audit log records every action by users, service accounts and the platform, at organization, folder and project scope, and is enabled by default rather than requiring configuration. That covers the authentication, authorization and administrative categories above without you building anything, which is a substantial head start.

The property to plan around is retention: 90 days in the Portal. For security purposes that is short relative to typical dwell time, and the mechanism for extending it is exporting through Telemetry Router to a destination you choose, with a worked example of connecting it to Logs.

That export is the step that turns a 90-day operational record into a security record, and it has to exist before the intrusion you will investigate. Sending it to a project with separate credentials is what gives it the different trust boundary this best practice asks for.

Application-level security signals go to Observability or Logs and depend on your instrumentation under OPS 7. Their retention has a ceiling: the Observability service plans default logs and traces to 5 days with a maximum of 30. That is a sensible operational figure and well short of typical intrusion dwell time, which means security retention is an export decision rather than a settings decision, in the same way the audit log is.

Tradeoffs. Cost Optimization. Long retention of security logs accumulates continuously and its value is invisible until an investigation, which is why it gets cut. OPS 7.4 is the argument for recording the reason next to the number.

Verify. If an attacker gained administrative access to your environment today, which records of that could they delete? How far back does your audit history actually reach?


SEC 11.2 Alert on patterns that indicate compromise rather than on volume

Section titled “SEC 11.2 Alert on patterns that indicate compromise rather than on volume”

Risk if not established: High

Collecting signals is necessary and produces nothing on its own. What turns collection into detection is a small set of alerts on patterns that genuinely indicate something wrong.

The patterns that earn an alert are the ones that are rare in normal operation and characteristic of an intrusion:

  • Authentication anomalies. A service account authenticating from somewhere new, a burst of failures followed by a success, use of a credential that has been dormant.
  • Privilege escalation. A role granted at organization scope, an identity granted Owner, a break-glass account used.
  • Unusual data access. Bulk reads of a classified data set, access outside working patterns.
  • Outbound anomalies. Traffic to destinations outside the allowlist, or volumes that do not fit the workload.

Keep the set small deliberately. The failure mode here is a hundred rules producing a hundred alerts a day, at which point nobody reads any of them and the detection capability is nominal. Two alerts a week that always mean something beat that comfortably.

Tune with the environment rather than against a template. What is anomalous depends on what is normal, and normal is specific to your workload.

On STACKIT. CSPM detects configuration weaknesses, which is posture rather than activity. It answers “is something exposed” and not “is someone using it”. Both are needed and they are different questions.

For activity, alerting is built on the audit records exported under SEC 11.1 and on application signals in Observability , using alert groups to separate security alerts from operational ones so that they route differently.

There is no managed security information and event management service, so correlation rules across signals are built with the tooling you choose. That is the usual division, and for many workloads the small set of high-value alerts described above is achievable without a dedicated platform.

Tradeoffs. Operational Excellence. Tuning is continuous, and an untuned rule set becomes noise within weeks. Reviewing which alerts fired and which led to action, as REL 10.2 suggests, is what keeps it honest.

Verify. Which security alerts exist, how many fired last month, and how many led to an investigation? If a service account key were used from an unexpected source tonight, what would happen?


SEC 11.3 Define the response before an alert fires

Section titled “SEC 11.3 Define the response before an alert fires”

Risk if not established: High

Security incidents share a structure with operational ones, which OPS 9.1 covers, and add requirements that change the response.

Evidence preservation comes before mitigation more often than in an operational incident. The instinct to terminate a compromised instance destroys the memory state, the process list and the attacker’s tooling. Snapshot first where the severity warrants it.

Containment is a distinct phase. Before eradicating, stop the spread: revoke the credential, isolate the network segment, disable the account. Doing this well depends on the segmentation from SEC 2 actually existing.

The infrastructure may not be trustworthy. If an attacker has administrative access, your monitoring, your deployment pipeline and your communication channels may all be observed. A response plan that assumes they are clean is a plan the attacker is reading.

Legal and regulatory duties have clocks. Personal data breaches carry notification deadlines that begin at awareness rather than at resolution, which means legal counsel is part of the response rather than something that follows it. SOV 8 establishes which obligations apply.

Name who decides. Isolating a production system to contain an intrusion is a costly decision, and it needs an owner who can make it quickly.

On STACKIT. The audit log is the primary investigation source for what an attacker did on the platform, subject to the retention and export considerations in SEC 11.1.

Containment mechanisms are the ones already covered: revoking a service account key or a role binding under access and identity , and network isolation through security groups or the Unified Firewall . Knowing which of those you would use, and having the permission to use it, belongs in the plan rather than in the incident.

status.stackit.cloud is where a platform-side event would appear, which is worth checking early to establish whether what you are seeing is yours.

Tradeoffs. Reliability. Containment actions cause outages by design. Isolating a system stops the spread and stops the service, and that trade is a decision somebody has to be authorized to make.

Verify. Who can revoke a compromised credential in your environment right now, and how long would it take? What is your notification deadline for a personal data breach, and who starts that clock?


SEC 11.4 Rehearse it, including the parts that are not technical

Section titled “SEC 11.4 Rehearse it, including the parts that are not technical”

Risk if not established: Medium

Security incident response is rehearsed less than any other procedure in this framework, because it is unpleasant and because the scenarios feel unlikely until they are not.

Rehearse at the depth the risk justifies. A tabletop exercise costs a few hours and finds most of the process gaps: who decides, who is called, what is said, and which permission nobody has. A technical exercise verifies that containment actually works and that the evidence you assumed exists does.

Rehearse the non-technical parts, since they are where the delays are. Notification duties, customer communication, and the question of who talks to a regulator are all decisions people make badly under pressure and well in advance.

Include the assumption that the environment is compromised. A rehearsal where the monitoring, the pipeline and the chat channel are all trusted is a rehearsal of the easy case.

Record what it finds and track it like any other incident action under OPS 9.4. A rehearsal whose findings are not closed produced a document rather than a capability.

On STACKIT. Rehearsal needs an environment, which OPS 6 and OPS 3 provide: creating a representative environment from definitions is what makes a technical exercise affordable without touching production.

Verify the evidence path during the exercise rather than assuming it. Whether audit records were actually exported through Telemetry Router , and whether anyone can query them, is the kind of thing that is configured once and never checked until it matters.

Tradeoffs. Cost Optimization. Rehearsal time from people whose time is expensive, for a scenario that may not occur. It is the same argument as REL 9.3 and it has the same answer: the alternative is a response capability nobody has evidence for.

Verify. When did you last rehearse a security incident, what scenario, and what did it find? Did the rehearsal include the notification and communication steps?


  • OPS 9 Incident management, whose process this extends
  • SEC 1 Security baseline, which detects configuration weakness rather than activity
  • SEC 2 Segmentation, which containment depends on
  • SEC 5.4 Entitlement review, which uses the same records
  • SOV 7 Auditability and SOV 8 Compliance mapping, which set the retention and the duties