Skip to content
Beta

Big Data Processing with Apache Spark

In 1 trail

Last updated on

A course about harnessing distributed computing for data at scale — and, just as important, about knowing when you actually need it: not every dataset requires Spark, sometimes PostgreSQL is all you need. It covers Spark’s architecture and execution model, then moves into writing efficient PySpark code, implementing stream processing pipelines, and optimizing jobs for performance and resource use.

Across four modules it goes from “what is Big Data and when do you need distributed processing” through PySpark DataFrame operations, structured streaming with Kafka sources, and performance tuning and production operations — deploying and monitoring Spark on STACKIT infrastructure, including Kubernetes.

  • Define “Big Data” and identify when distributed processing is necessary
  • Understand Spark’s architecture and execution model
  • Write efficient PySpark code for data transformation and analysis
  • Implement stream processing pipelines for continuous data flows
  • Optimize Spark jobs for performance and resource utilization
  • Choose between Spark and traditional databases based on data characteristics
  • Deploy and monitor Spark applications on STACKIT infrastructure
  • Understanding Big Data and Apache Spark
  • Data Processing with PySpark
  • Stream Processing with Structured Streaming
  • Performance Tuning and Production Operations
STACKIT documentation university.stackit.cloud View Course on STACKIT University Open the documentation
Asset historyActive 2 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
Maintainers
TMTobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172CC.C1SCF Core · STACKITOwnerActive 5 of the last 12 weeks · 23 updatesSTACKITwww.linkedin.com/in/can-celik-645932315can.celik1@digits.schwarzContributed in STACKIT