---
title: "Kubernetes for Data Engineering"
description: "Run data workloads on Kubernetes: stateful services and persistent storage, Airflow and Spark on Kubernetes, Kafka, and data platform operations."
scfAsset:
  managed: false
  category: "guide"
  maintainers:
    - user: "can.celik1"
  external: true
  tags: ["STACKIT University", "Learning", "Kubernetes", "Data Engineering"]
source_url: "https://framework.stackit.cloud/data-and-ai/assetcontainer/stackit/kubernetes-data-engineering/"
source_file: "docs/data-and-ai/assetcontainer/stackit/kubernetes-data-engineering.mdx"
---

## Course Overview

Kubernetes excels at stateless applications, but data engineering brings different challenges: stateful workloads, persistent storage, ordered operations, and resource-intensive jobs. This course, building on basic Kubernetes knowledge, addresses those challenges head-on and teaches production-ready data infrastructure on Kubernetes.

Across four modules it covers StatefulSets and persistent storage for databases, running Apache Airflow on Kubernetes with the KubernetesExecutor, deploying distributed processing frameworks (Spark, Flink) and Kafka, and data platform operations — resource management, monitoring, security, and disaster recovery.

### What You'll Learn
- Deploy and manage stateful data services on Kubernetes
- Implement persistent storage strategies for databases and data lakes
- Run Apache Airflow on Kubernetes with the KubernetesExecutor
- Deploy distributed data processing frameworks (Spark, Flink)
- Build event streaming platforms with Kafka on Kubernetes
- Optimize resource allocation for data-intensive workloads
- Handle backup, disaster recovery, and data migration

### Modules
- Stateful Workloads and Storage
- Airflow on Kubernetes
- Distributed Processing Frameworks
- Data Platform Operations

<LinkCard title="View Course on STACKIT University" href="https://university.stackit.cloud/totara/catalog/index.php" />
