---
title: "Architecture Asset: Generative AI with Retrieval-Augmented Generation on STACKIT"
description: 'Reference architecture for a sovereign RAG application on STACKIT: documents in Object Storage, embeddings on Kubernetes, OpenSearch, and model serving.'
sidebar:
  badge:
    text: "STACKIT"
    variant: success
scfAsset:
  maintainers:
    - user: "tobias.mueller"
  managed: false
  category: 'blueprint'
  external: false
  tags: ["build", "generative-ai", "rag", "llm", "opensearch", "wip"]
source_url: "https://framework.stackit.cloud/data-and-ai/assetcontainer/stackit/generative-ai-rag/"
source_file: "docs/data-and-ai/assetcontainer/stackit/generative-ai-rag.mdx"
---

## Overview

This pattern delivers grounded generative AI. Documents are embedded into an OpenSearch vector store; at query time the app retrieves relevant context and calls a served LLM, keeping data and inference on sovereign infrastructure.

## Typical use case

- **Knowledge assistants**: answer questions grounded in internal documents.
- **Support automation**: draft responses from trusted knowledge sources.
- **Sovereign GenAI**: run retrieval and inference with control over data and residency.

## Architecture diagram

```d2
vars: {
  d2-config: {
    pad: 32
  }
}

style.font-size: 22
direction: down
grid-columns: 1

Ingestion: "Knowledge Ingestion" {
  direction: right
  grid-columns: 3

  Docs: "Documents (Object Storage)" {
    icon: ../../../../../../public/stackit-icons/computing/object-storage.svg
    link: https://docs.stackit.cloud/products/storage/object-storage/
  }
  Embed: "Embedding Pipeline (Kubernetes)" {
    icon: ../../../../../../public/stackit-icons/artificial-intelligence/workflows.svg
    link: https://docs.stackit.cloud/products/runtime/kubernetes-engine/
  }
  Vector: "Vector Store (OpenSearch)" {
    icon: ../../../../../../public/stackit-icons/databases/opensearch.svg
    link: https://docs.stackit.cloud/products/databases/opensearch/
  }
}

Serving: "Query & Generation" {
  direction: right
  grid-columns: 1

  LB: "API Gateway / LB" {
    icon: ../../../../../../public/stackit-icons/networking/application-load-balancer.svg
    link: https://docs.stackit.cloud/products/network/load-balancing-and-content-delivery/application-load-balancer/
  }

  App: "RAG Application (Kubernetes)" {
    icon: ../../../../../../public/stackit-icons/runtime/kubernetes.svg
    link: https://docs.stackit.cloud/products/runtime/kubernetes-engine/
  }

  LLM: "Model Serving (LLM)" {
    icon: ../../../../../../public/stackit-icons/artificial-intelligence/model-serving.svg
    link: https://docs.stackit.cloud/products/data-and-ai/ai-model-serving/
  }
}

Cross: "Platform Services" {
  direction: right
  grid-columns: 2

  Secrets: "Secret Manager" {
    icon: ../../../../../../public/stackit-icons/security/secrets-manager.svg
    link: https://docs.stackit.cloud/products/security/secrets-manager/
  }
  Obs: "Observability" {
    icon: ../../../../../../public/stackit-icons/logging-monitoring/observability.svg
    link: https://docs.stackit.cloud/products/logging-and-monitoring/observability/
  }
}

Ingestion.Docs -> Ingestion.Embed
Ingestion.Embed -> Ingestion.Vector
Serving.LB -> Serving.App
Serving.App -> Ingestion.Vector
Serving.App -> Serving.LLM
Serving -> Cross
```

## Design best practices

- **Ground every answer**: retrieve from the vector store before generation to reduce hallucination.
- **Keep data sovereign**: run embedding, retrieval, and inference on STACKIT.
- **Evaluate continuously**: track answer quality, safety, and cost before and after release.
- **Secure prompts and secrets**: isolate keys and never log sensitive context.
