Skip to content
Beta

STACKIT Intake

In 1 trail

Last updated on

STACKIT Intake simplifies real-time data ingestion by accepting Apache Kafka protocol streams and automatically writing them to Apache Iceberg lakehouse tables.

  • Automated Lakehouse Streaming: Ingests high-volume streams into structured Iceberg tables without manual ETL code.
  • Protocol Compatibility: Uses standard Apache Kafka interfaces, allowing existing producers to connect seamlessly.
  • Resilient Ingestion: Buffers incoming data up to 24 hours during downstream target processing outages.
  • Architecture: Managed runner infrastructure with auto-scaling capabilities.
  • Lakehouse Bridge: Automatically syncs incoming streams into the STACKIT Dremio catalog and Iceberg storage.
  • Security: Secured via technical user credentials and Personal Access Tokens (PAT).

The values below come from the STACKIT documentation and update themselves.

From the STACKIT docsArchitecture › IntakeSource updated 17.03.2026 · copied 05.10.2026

An Intake represents a complete data ingestion pipeline into a single Apache Iceberg table. It ingests data in JSON format and periodically flushes it to the corresponding Dremio table, ensuring each message is inserted exactly once. Flushing introduces a delay of around 5 minutes before messages appear in the table.

An Intake can either write to an existing table or dynamically create one, inferring the schema from the first message it receives. To handle potential downstream outages, such as Dremio maintenance, Intakes buffer messages for up to 24 hours.

An Intake is defined by:

  • Dremio Instance Parameters & User Credentials: URLs for the Iceberg REST Catalog and authentication endpoint, along with a Dremio Personal Access Token (PAT) for user authentication.
  • Iceberg Table Definition: The namespace and name of the target table. If not provided, a default namespace and an auto-generated name are used.
  • Partitioning Information: The option to partition the Iceberg table based on a field in the JSON message or automatically by ingestion date using the automatically added __intake_ts column. If no partition option is given, an Intake will not create any partitions in the result table.

Each Intake provides two Kafka protocol topics for communication:

  • Intake Topic: The primary topic for sending messages to be ingested.
  • Dead Letter Queue (DLQ) Topic: A topic where undeliverable messages (e.g., non-JSON messages) are temporarily stored for debugging purposes.
What is this?

This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.

  • Fixed Sink Target: Currently supports STACKIT Dremio (Apache Iceberg tables) as the sole sink destination.
  • JSON Payload Requirement: Schema derivation requires incoming stream messages to be valid JSON structures.
  • Flush Latency: Features an intentional micro-batching flush delay of approximately 5 minutes before data is visible in tables.
STACKIT documentation docs.stackit.cloud STACKIT Intake Overview Open the documentation
Asset historyActive 2 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
Maintainers
  • ?Name not public?Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site.
TMTobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172??Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site.Contributed in STACKIT
  • Tobias M.Tobias M.Head of STACKIT Cloud Framework · STACKITOwnerActive 12 of the last 12 weeks · 168 updatesSTACKITwww.linkedin.com/in/tobias-müller-011304172 · Oct 5, 2026

  • Name not public?Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site. · Sep 9, 2026