Skip to content
Beta

Kubernetes for Data Engineering

In 1 trail

Last updated on

Kubernetes excels at stateless applications, but data engineering brings different challenges: stateful workloads, persistent storage, ordered operations, and resource-intensive jobs. This course, building on basic Kubernetes knowledge, addresses those challenges head-on and teaches production-ready data infrastructure on Kubernetes.

Across four modules it covers StatefulSets and persistent storage for databases, running Apache Airflow on Kubernetes with the KubernetesExecutor, deploying distributed processing frameworks (Spark, Flink) and Kafka, and data platform operations — resource management, monitoring, security, and disaster recovery.

  • Deploy and manage stateful data services on Kubernetes
  • Implement persistent storage strategies for databases and data lakes
  • Run Apache Airflow on Kubernetes with the KubernetesExecutor
  • Deploy distributed data processing frameworks (Spark, Flink)
  • Build event streaming platforms with Kafka on Kubernetes
  • Optimize resource allocation for data-intensive workloads
  • Handle backup, disaster recovery, and data migration
  • Stateful Workloads and Storage
  • Airflow on Kubernetes
  • Distributed Processing Frameworks
  • Data Platform Operations
STACKIT documentation university.stackit.cloud View Course on STACKIT University Open the documentation
Asset historyActive 2 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
Maintainers
TMTobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172CC.C1SCF Core · STACKITOwnerActive 5 of the last 12 weeks · 23 updatesSTACKITwww.linkedin.com/in/can-celik-645932315can.celik1@digits.schwarzContributed in STACKIT