Course Overview
Section titled “Course Overview”A course about harnessing distributed computing for data at scale — and, just as important, about knowing when you actually need it: not every dataset requires Spark, sometimes PostgreSQL is all you need. It covers Spark’s architecture and execution model, then moves into writing efficient PySpark code, implementing stream processing pipelines, and optimizing jobs for performance and resource use.
Across four modules it goes from “what is Big Data and when do you need distributed processing” through PySpark DataFrame operations, structured streaming with Kafka sources, and performance tuning and production operations — deploying and monitoring Spark on STACKIT infrastructure, including Kubernetes.
What You’ll Learn
Section titled “What You’ll Learn”- Define “Big Data” and identify when distributed processing is necessary
- Understand Spark’s architecture and execution model
- Write efficient PySpark code for data transformation and analysis
- Implement stream processing pipelines for continuous data flows
- Optimize Spark jobs for performance and resource utilization
- Choose between Spark and traditional databases based on data characteristics
- Deploy and monitor Spark applications on STACKIT infrastructure
Modules
Section titled “Modules”- Understanding Big Data and Apache Spark
- Data Processing with PySpark
- Stream Processing with Structured Streaming
- Performance Tuning and Production Operations
Asset historyActive 2 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
- CCC.C1SCF Core · STACKITOwner
C.C1SCF Core · STACKITOwnerActive 5 of the last 12 weeks · 23 updateswww.linkedin.com/in/can-celik-645932315can.celik1@digits.schwarz