---
title: "Data Lakehouse with Dremio (DeepDive)"
description: "Take a technical deep dive into Dremio on STACKIT: Apache Arrow, query federation, the semantic layer, Data Reflections, Apache Iceberg, and Project Nessie."
scfAsset:
  managed: false
  category: "guide"
  maintainers:
    - user: "can.celik1"
  external: true
  tags: ["STACKIT University", "Learning", "Dremio", "Apache Iceberg", "Apache Arrow", "Nessie"]
source_url: "https://framework.stackit.cloud/data-and-ai/assetcontainer/stackit/data-lakehouse-dremio/"
source_file: "docs/data-and-ai/assetcontainer/stackit/data-lakehouse-dremio.mdx"
---

## Course Overview

The scenario this course is built around: your organization needs a unified analytics platform where business analysts get fast, self-service access to curated data products while data scientists query raw data in place — without moving or copying it unnecessarily, and without breaking GDPR compliance along the way. You'll start with the strategic case for the STACKIT and Dremio partnership, then go deep on the technical internals: Dremio's distributed query engine, how Apache Arrow enables zero-copy, vectorized execution, and how query federation lets Dremio analyze data across multiple sources without ever moving it.

From there the course builds the layers that make a lakehouse usable at organizational scale — a Universal Semantic Layer for consistent, business-friendly data products, Data Reflections to accelerate query performance dramatically without duplicating pipelines, and Apache Iceberg with Project Nessie for ACID guarantees, schema evolution, time travel, and Git-like branching over your data. Each of the three technical modules pairs the concept with a hands-on lab, so you leave able to both explain the architecture and actually operate it.

### What You'll Learn
- Explain how Dremio's distributed engine and Apache Arrow deliver zero-copy, vectorized query performance
- Run federated queries across multiple data sources without moving or duplicating data
- Build a semantic layer and use Data Reflections to accelerate query performance
- Apply Apache Iceberg and Project Nessie for ACID tables, schema evolution, and Git-style data branching

### Modules
- Dremio Architecture and Apache Arrow
- Query Federation and Semantic Layer
- Reflections and Performance
- Apache Iceberg and Project Nessie

<LinkCard title="View Course on STACKIT University" href="https://university.stackit.cloud/totara/catalog/index.php" />
