Cloud Computing Fundamentals
STACKIT
No login. Sent to the Framework Core Team. The page link is included automatically.
Send your feedback via email now.
Open your email client and create a new email.
Copy the email address and paste it into the To: field:
Copy your feedback text and paste it into the email body (Ctrl+V / Cmd+V):
Last updated on
Comprehensive learning path for data engineers: SQL, dimensional modeling, pipelines with Airflow and Spark, and data lakehouses with Iceberg and Dremio on STACKIT.
STACKIT
Build a solid foundation in cloud computing: service and deployment models, core architecture patterns, and security principles for sovereign environments.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
A comprehensive introduction to cloud computing for anyone new to the field, structured around three building blocks: what “the cloud” actually is and how it differs from traditional IT, the service and deployment models providers offer (IaaS, PaaS, SaaS; public, private, hybrid, and multi-cloud), and the architectural and security principles that hold a cloud environment together. Each module is a self-contained, beginner-friendly introduction — no prior cloud experience assumed.
This is the entry point for the Data Engineer path: before working with STACKIT’s data and lakehouse services, it’s worth having a precise, shared vocabulary and mental model for what cloud computing is and isn’t.
STACKIT
Get hands-on with the STACKIT Portal, core IaaS and PaaS services, and the project and resource model behind every STACKIT deployment.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
The foundational entry point into the STACKIT sovereign cloud, structured as five progressively hands-on modules. You start with an orientation to STACKIT itself, then move into the STACKIT Portal and management interfaces where day-to-day work actually happens, before working through the core infrastructure (IaaS) services and the platform and application (PaaS) services that sit on top of them. The course closes by covering how projects, resources, and organizations are structured — the account model every subsequent action on the platform depends on.
This is the first STACKIT-specific course in every learning path, and deliberately so: before you provision a VM, deploy an app, or configure networking, you need to know where things live in the platform’s project/organization hierarchy and which interface — Portal, CLI, API, or Terraform — fits the task at hand. Getting this structure right early avoids the misconfigured projects and permission confusion that otherwise surface later, once real workloads depend on them.
STACKIT
Learn the fundamentals of data engineering: the data engineer's role, the data lifecycle, modern data architectures, and governance and security.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
This foundational course introduces the role and responsibilities of a data engineer and maps out the complete data lifecycle, from generation to analysis, before tracing how data architectures have evolved — from traditional data warehouses, through data lakes, to the lakehouse and data mesh patterns used in modern platforms. It’s the bedrock course of the Data Engineer learning path: everything that follows, from Airflow orchestration to Dremio lakehouses, assumes the shared vocabulary and mental model built here.
The course closes on governance and security, treating them not as an afterthought but as core non-functional requirements of any data platform — particularly relevant in a sovereign cloud context where data residency and access control carry real regulatory weight. If you’re new to data engineering or coming from an adjacent role like software development or analytics, this is the course that establishes why the discipline exists and what a data engineer is actually responsible for day to day.
STACKIT
Relational database concepts, SQL querying, query optimization, and basic database administration as a foundation for data engineering.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
A foundational course in relational databases and SQL for anyone about to work with data warehouses and pipelines. It starts with relational database concepts — tables, keys, and normalization — before building up SQL querying skills from basic selects and joins to more advanced querying and optimization, and closes with basic database administration.
It’s a prerequisite for the Data Modeling and Warehousing course: comfort with joins, normalization, and basic database design is assumed from here on.
STACKIT
Design dimensional models for analytics: star and snowflake schemas, fact and dimension tables, slowly changing dimensions, and warehouse architecture.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
Data modeling is what separates good data engineers from great ones: this course teaches the art and science of designing data structures optimized for analytics and reporting, transforming normalized transactional data into dimensional models that enable fast business intelligence. It covers both the “how” and the “why” behind dimensional modeling — when to denormalize for performance and how to handle scenarios like slowly changing dimensions.
Across four modules it moves from dimensional modeling foundations (facts, dimensions, star schemas) to building fact and dimension tables, handling slowly changing dimensions with the standard SCD strategies, and warehouse architecture patterns (Kimball vs. Inmon) with performance optimization.
STACKIT
Learn the core concepts of the data lakehouse architecture: open table formats, how it compares to data warehouses and data lakes, and lakehouse strategy.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
This course introduces the data lakehouse architecture and the open table formats that underpin it, working through what a lakehouse actually is, how it differs architecturally from a traditional data warehouse or a plain data lake, and where each approach still makes sense on its own. It’s a foundational, vendor-agnostic course: the goal is to leave you able to reason about the trade-offs between these architectures before you commit to a specific implementation.
That grounding matters because it’s a prerequisite for the more specialized data engineering courses that follow, including the Apache Iceberg and Dremio deep dives — this is where you build the conceptual map before working with a specific table format or query engine. The closing module turns theory into decision-making, walking through how to evaluate and build a lakehouse strategy for your own organization rather than adopting one by default.
STACKIT
Apache Iceberg for the data lakehouse: table architecture, table management, schema evolution and partitioning, and time travel and branching.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
An open table format course covering Apache Iceberg, the technology behind modern data lakehouse architectures. It builds on Data Lakehouse Fundamentals with a deep dive into how Iceberg tables are structured and managed: table operations, schema evolution, partitioning strategies, and the time-travel and branching features that make lakehouse data behave a lot like a version-controlled dataset.
STACKIT
Learn to orchestrate data pipelines with Apache Airflow: DAGs, operators and sensors, Kubernetes integration, and CI/CD for production workflows on STACKIT.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
This course teaches you to design, build, and operate production-grade data pipelines with Apache Airflow. The challenge in modern data engineering usually isn’t processing data — it’s coordinating hundreds of interdependent tasks, handling failures gracefully, and keeping visibility into what actually ran. You’ll start with Airflow’s core architecture (scheduler, executor, workers, metadata database) and DAG fundamentals, then build up through operators, sensors, and common ETL/ELT patterns, including data quality gates and incremental loading.
The second half is where the course earns its “production-grade” framing: dynamic DAG generation, cross-DAG dependencies with external task sensors, retry and backfill strategies, and — critically for STACKIT — running Airflow on STACKIT Kubernetes Engine with the KubernetesPodOperator for task isolation, plus the CI/CD practices needed to ship DAGs safely. Every module pairs the concept with a practical exercise deployed against real STACKIT infrastructure, so the skills transfer directly to pipelines you’d actually operate.
STACKIT
Distributed data processing with Apache Spark: PySpark DataFrames, structured streaming, performance tuning, and production deployment on STACKIT.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
A course about harnessing distributed computing for data at scale — and, just as important, about knowing when you actually need it: not every dataset requires Spark, sometimes PostgreSQL is all you need. It covers Spark’s architecture and execution model, then moves into writing efficient PySpark code, implementing stream processing pipelines, and optimizing jobs for performance and resource use.
Across four modules it goes from “what is Big Data and when do you need distributed processing” through PySpark DataFrame operations, structured streaming with Kafka sources, and performance tuning and production operations — deploying and monitoring Spark on STACKIT infrastructure, including Kubernetes.
STACKIT
Take a technical deep dive into Dremio on STACKIT: Apache Arrow, query federation, the semantic layer, Data Reflections, Apache Iceberg, and Project Nessie.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
The scenario this course is built around: your organization needs a unified analytics platform where business analysts get fast, self-service access to curated data products while data scientists query raw data in place — without moving or copying it unnecessarily, and without breaking GDPR compliance along the way. You’ll start with the strategic case for the STACKIT and Dremio partnership, then go deep on the technical internals: Dremio’s distributed query engine, how Apache Arrow enables zero-copy, vectorized execution, and how query federation lets Dremio analyze data across multiple sources without ever moving it.
From there the course builds the layers that make a lakehouse usable at organizational scale — a Universal Semantic Layer for consistent, business-friendly data products, Data Reflections to accelerate query performance dramatically without duplicating pipelines, and Apache Iceberg with Project Nessie for ACID guarantees, schema evolution, time travel, and Git-like branching over your data. Each of the three technical modules pairs the concept with a hands-on lab, so you leave able to both explain the architecture and actually operate it.
STACKIT
A comprehensive introduction to Docker and containerization: building images, running containers, and orchestrating them with Docker Compose.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
A hands-on introduction to containerization with Docker: how it packages applications and their dependencies into isolated, portable containers so they run the same on a laptop, in a development environment, or in the cloud. The course walks through building and managing containers efficiently, then orchestrating multi-container applications with Docker Compose, using a “Hello STACKIT” website as a running practical exercise.
It’s built as the foundation for the follow-on Kubernetes course, both on STACKIT VMs and locally with Docker Desktop.
STACKIT
Master Kubernetes architecture, core objects, networking, and storage, then apply it hands-on with the STACKIT Kubernetes Engine (SKE).
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
An eight-module, hands-on introduction to Kubernetes built for cloud engineers approaching the topic for the first time — no prior Kubernetes knowledge required, just basic Docker and Linux familiarity. You start with cluster architecture and a local environment (Minikube or Kind), then work through core objects and YAML manifests, Services and Ingress, persistent storage and StatefulSets, RBAC and secrets management, and scheduling and scaling, building fluency with kubectl and multi-container patterns like sidecars and init containers along the way.
The course can be completed entirely in a local cluster or on a live STACKIT Kubernetes Engine (SKE) cluster, and a dedicated module walks through deploying your own application to SKE so the concepts land as production practice, not just theory. A closing module leaves you with links, tips, and exam guidance for the CNCF CKAD certification, making this course a solid foundation whether your goal is day-to-day cluster operations or a formal certification path.
STACKIT
Run data workloads on Kubernetes: stateful services and persistent storage, Airflow and Spark on Kubernetes, Kafka, and data platform operations.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
Kubernetes excels at stateless applications, but data engineering brings different challenges: stateful workloads, persistent storage, ordered operations, and resource-intensive jobs. This course, building on basic Kubernetes knowledge, addresses those challenges head-on and teaches production-ready data infrastructure on Kubernetes.
Across four modules it covers StatefulSets and persistent storage for databases, running Apache Airflow on Kubernetes with the KubernetesExecutor, deploying distributed processing frameworks (Spark, Flink) and Kafka, and data platform operations — resource management, monitoring, security, and disaster recovery.
STACKIT
Get started with STACKIT Dremio, the managed sovereign data lakehouse: instance setup, SSO/OIDC authentication, and connecting data sources.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
An introduction to STACKIT Dremio, a fully managed data lakehouse service running entirely on STACKIT’s sovereign infrastructure — Kubernetes orchestration, elastic engine nodes, and storage integration are all handled for you. The course covers what sets STACKIT Dremio apart from Dremio Cloud and self-hosted Dremio, then walks through creating and configuring your own instance, setting up SSO/OIDC authentication and access policies, and connecting data sources through the Open Catalog.
It’s the practical bridge between the conceptual Data Lakehouse and Dremio (DeepDive) courses and the guided exercises in STACKIT Dremio: Hands-On Lab.
STACKIT
Build a complete data lakehouse on STACKIT Dremio in this hands-on lab: instance setup, Iceberg tables, the semantic layer, Data Reflections, and Nessie branching.
The sovereign European cloud provider behind the framework, delivering IaaS and PaaS from German and Austrian data centers with full digital independence.
This is a fully hands-on lab, not a conceptual course: each of its four modules is a self-contained lab that builds directly on the last, taking you from an empty STACKIT project to a working lakehouse. You’ll provision a STACKIT Dremio instance, connect Object Storage, upload sample CSV and Parquet files, and convert them into Iceberg tables you can query with SQL — the same foundational steps you’d follow in a real deployment.
From there you build out what makes the lakehouse useful in practice: a semantic layer with business-friendly virtual datasets, Raw and Aggregation Reflections to compare accelerated versus unaccelerated query performance side by side, and in the final lab, a second federated data source (PostgreSQL), row-level security, data masking, and Nessie branching so you can experiment on your data without risk. It’s the direct hands-on companion to the Dremio DeepDive and Getting Started courses — where those explain the concepts, this is where you build them yourself.
Tobias M.Tobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172
· Aug 26, 2026
Tobias M.Tobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172
· Aug 19, 2026
C.C1C.C1SCF Core · STACKITOwnerActive 5 of the last 12 weeks · 23 updatesSTACKITwww.linkedin.com/in/can-celik-645932315can.celik1@digits.schwarz
· Aug 10, 2026
C.C1C.C1SCF Core · STACKITOwnerActive 5 of the last 12 weeks · 23 updatesSTACKITwww.linkedin.com/in/can-celik-645932315can.celik1@digits.schwarz
· Aug 10, 2026
C.C1C.C1SCF Core · STACKITOwnerActive 5 of the last 12 weeks · 23 updatesSTACKITwww.linkedin.com/in/can-celik-645932315can.celik1@digits.schwarz
· Aug 4, 2026
External link
This link goes to an external site outside STACKIT. We do not vet third-party content or downloads, so follow it only if you trust the source.