Skip to content
Beta

Service Management and Operations

Last updated on

Service management makes cloud operations repeatable after a migration wave. It aligns workload teams, the cloud platform team, and service stakeholders around how services are monitored, supported, changed, and restored.

  • Incident management: Define severity levels, service ownership, communication channels, escalation paths, and post-incident reviews.
  • Change management: Apply risk-based change controls that preserve delivery speed while protecting shared STACKIT landing-zone components and production workloads.
  • Problem management: Track recurring faults, configuration drift, and systemic causes through a prioritized improvement backlog.
  • Request management: Define fulfilment paths for access, project onboarding, platform services, and approved exceptions.

Before a workload leaves migration hypercare, confirm that the responsible team has:

  • Monitoring, alert routing, logs, and dashboards aligned with its service targets.
  • Documented backup, recovery, and rollback procedures that reflect the selected STACKIT services.
  • A tested support path, including customer support-plan escalation where applicable.
  • Runbooks, ownership data, dependencies, and known-risk records available to operators.
  • Approved change and release procedures for the workload and its shared platform dependencies.

Define a small set of service level objectives and error budgets for critical customer outcomes. Use them in service reviews together with security findings, operational incidents, and cost trends. Targets should shape priorities; they are not a substitute for understanding application dependencies and recovery requirements.

Asset title
Framework
Asset type