Availability SLA
Monthly availability of the service. Based on the STACKIT platform SLA, reduced by an internal operational buffer. Transparently communicated and documented in the service catalogue.
ITIL (IT Infrastructure Library) was the response to the chaos of early IT organisations: standardised processes for change management, incident management, problem management and service design. For most organisations, ITIL represented a necessary step towards maturity.
In cloud environments, ITIL comes under pressure. A Change Advisory Board that meets once a week is not compatible with an organisation aiming for hundreds of deployments per day. A 14-day change request process blocks the agile iteration that the cloud is designed to enable.
The answer is not: abolish ITIL. The answer is: evolve ITIL.
What remains from ITIL:
What changes fundamentally:
Classic change management was designed for a world in which every change to production systems was manual, risky, and difficult to reverse. In that world, a formal approval process made sense.
In cloud environments with Infrastructure-as-Code and automated tests, the risk assessment changes fundamentally. A Terraform change that has been automatically checked against policies before deployment, tested in staging, and reviewed by two people via pull request has a different risk profile than a manual configuration change at night.
The modernised change model:
Standard Changes (frequent, well-documented, low risk) are fully automated. No CAB, no change ticket — the automation is the approval process. Examples: container image updates, configuration changes for known parameters, resource scaling.
Normal Changes (changes to critical system components) go through a pull request process with two reviewers, an automated test suite, and a deployment window. The “CAB” is the asynchronous review by experienced colleagues — faster, more efficient, with the same level of assurance.
Emergency Changes (critical production fixes) have an accelerated process with documentation completed afterwards. They are discussed in the post-incident review: was the emergency procedure correctly applied? What prevents the same issue from being treated as an emergency again?
The core model of incident management — detection, triage, escalation, resolution, post-mortem — is just as valid in cloud environments as in classical IT environments.
What changes is the speed and the expectations.
Detection: Classically through user reports or manual monitoring. Today through automated observability systems that detect anomalies before users are affected.
Triage: Classically through manual diagnosis via SSH and log files. Today through central dashboards, distributed tracing and structured log search.
On-call structure: Cloud environments require a 24/7 on-call rotation. This is culturally the biggest shift for many IT organisations: who is on call, how is it compensated, how is burnout prevented?
Post-mortem culture: Blameless post-mortems are standard in DevOps cultures — no blame assignment, but systemic learning. For ITIL-shaped organisations, this is often a cultural shift that requires time and leadership commitment.
The Configuration Management Database (CMDB) was ITIL’s answer to the question: what is running where, in what configuration, with what dependencies on what else? In classical environments it was valuable — and notoriously difficult to keep current.
In cloud environments with Infrastructure-as-Code, the CMDB is no longer a primary system. Infrastructure-as-Code (Terraform, Ansible) is the new CMDB — with the decisive advantage that it is automatically correct: what is in Terraform is (after the last apply) reality.
What this means for the organisation:
SLAs with business units remain important — but their design changes. Classical SLAs spoke about availability in percentages on a monthly basis. Cloud SLAs can be more granular.
Availability SLA
Monthly availability of the service. Based on the STACKIT platform SLA, reduced by an internal operational buffer. Transparently communicated and documented in the service catalogue.
Response Time SLA
Response times by severity. P1 within 15 minutes, P2 within 1 hour — not as a promise, but as a measurable commitment with monthly reporting.
Recovery Time SLA
RTO and RPO per tier class (Tier 1 to 4). Not as a theoretical figure, but as a regularly tested and demonstrated capability.
Change Lead Time
How long does it take to get an approved change into production? For standard changes: hours. For normal changes: days. Not weeks.
Abolishing ITIL processes overnight is just as wrong as transferring them unchanged into the cloud. The right approach is gradual.
Inventory: Which ITIL processes exist? Which of them make sense in the cloud, and which create drag?
Identify quick wins: Automating standard changes is typically fast to implement and immediately noticeable.
Revise the change process: Increase CAB frequency or replace with asynchronous review. Introduce an automation gateway as an alternative to manual approvals.
Introduce blameless post-mortems: Culturally demanding, but decisive for continuous improvement. Leadership must model that mistakes are learning opportunities.
Adapt CMDB strategy: Establish IaC as the primary configuration source; reduce the CMDB to strategic information.