---
title: "STACKIT AI Model Serving"
description: "Sovereign, managed API platform providing scalable inference for open-source Large Language Models (LLMs) via OpenAI-compatible interfaces."
scfAsset:
  category: "service"
  managed: true
  marketplaceUrl: "https://marketplace.stackit.cloud/en/products"
  tags: ["AI", "LLM", "Data and AI", "OpenAI", "Inference"]
  maintainers:
    - user: "alexander.gabert"
source_url: "https://framework.stackit.cloud/data-and-ai/assetcontainer/stackit/stackit-service-ai-model-serving/"
source_file: "docs/data-and-ai/assetcontainer/stackit/stackit-service-ai-model-serving.mdx"
---

STACKIT AI Model Serving gives developers turnkey access to hosted open-source Large Language Models (e.g., Llama 3, Mistral, Qwen, Gemma) using OpenAI-standard REST APIs.

## Service Overview
- **Sovereign GenAI**: Fully European-hosted AI execution keeping sensitive prompt data private and GDPR-compliant.
- **Cost Efficiency**: Pay-per-token pricing models eliminate the high cost of maintaining idle GPU hardware.
- **Drop-In Compatibility**: Zero code refactoring for applications built on OpenAI API structures.

## Technical Details
- **Deployment Options**: Shared Model Serving (cost-optimized, multi-tenant) and Custom Model Serving (dedicated, predictable performance).
- **Managed Scaling**: Auto-scaling inference infrastructure managing underlying NVIDIA GPU resources natively.

## Available Models

The list below comes from the STACKIT documentation and updates itself.

> From the STACKIT docs: [Available Shared Models of STACKIT AI Model Serving](https://docs.stackit.cloud/products/data-and-ai/ai-model-serving/basics/available-shared-models/) (Source updated 08.09.2026, copied 02.10.2026)

| Model | Capabilities | Context | Max. output | Status |
| --- | --- | --- | --- | --- |
| Qwen3-VL 235B | Chat, Vision | 200K | 16,384 | Supported |
| Qwen3.8 27B | Chat, Vision, Reasoning | 262K | 16,384 | Supported |
| Qwen3.6 27B | Chat, Vision, Reasoning | 262K | 16,384 | Deprecated |
| Llama 3.3 70B | Chat | 128K | 4,096 | Supported |
| GPT-OSS 120B | Chat, Reasoning | 131K | 8,192 | Supported |
| Gemma 3 27B | Chat, Vision | 131K | 4,096 | Deprecated |
| Gemma 4 31B | Chat, Vision, Reasoning | 256K | 4,096 | Supported |
| GPT-OSS 20B | Chat, Reasoning | 131K | 8,192 | Supported |
| E5 Mistral 7B | Embedding | n/a | n/a | Supported |
| Qwen3 Vision-Language Embedding | Embedding | n/a | n/a | Supported |

## Limitations & Constraints
- **Coarse Token Permissions**: API tokens grant access to all available shared models; fine-grained model-level access is unsupported.
- **Stateless Endpoint**: The API stores no session state or conversation context; context history must be sent with each request.
- **Shared Cluster Variance**: Shared tier latency fluctuates based on multi-tenant cluster demand and strict RPM/TPM rate limits.

<LinkCard title="STACKIT AI Model Serving Documentation" href="https://docs.stackit.cloud" />
