STACKIT AI Model Serving gives developers turnkey access to hosted open-source Large Language Models (e.g., Llama 3, Mistral, Qwen, Gemma) using OpenAI-standard REST APIs.
Service Overview
Section titled “Service Overview”- Sovereign GenAI: Fully European-hosted AI execution keeping sensitive prompt data private and GDPR-compliant.
- Cost Efficiency: Pay-per-token pricing models eliminate the high cost of maintaining idle GPU hardware.
- Drop-In Compatibility: Zero code refactoring for applications built on OpenAI API structures.
Technical Details
Section titled “Technical Details”- Deployment Options: Shared Model Serving (cost-optimized, multi-tenant) and Custom Model Serving (dedicated, predictable performance).
- Managed Scaling: Auto-scaling inference infrastructure managing underlying NVIDIA GPU resources natively.
Available Models
Section titled “Available Models”The list below comes from the STACKIT documentation and updates itself.
| Model | Capabilities | Context | Max. output | Status |
|---|---|---|---|---|
| Text models | ||||
| Qwen3-VL 235B | Chat, Vision | 200K | 16,384 | Supported |
| Qwen3.8 27B | Chat, Vision, Reasoning | 262K | 16,384 | Supported |
| Qwen3.6 27B | Chat, Vision, Reasoning | 262K | 16,384 | Deprecated |
| Llama 3.3 70B | Chat | 128K | 4,096 | Supported |
| GPT-OSS 120B | Chat, Reasoning | 131K | 8,192 | Supported |
| Gemma 3 27B | Chat, Vision | 131K | 4,096 | Deprecated |
| Gemma 4 31B | Chat, Vision, Reasoning | 256K | 4,096 | Supported |
| GPT-OSS 20B | Chat, Reasoning | 131K | 8,192 | Supported |
| Embedding models | ||||
| E5 Mistral 7B | Embedding | n/a | n/a | Supported |
| Qwen3 Vision-Language Embedding | Embedding | n/a | n/a | Supported |
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Limitations & Constraints
Section titled “Limitations & Constraints”- Coarse Token Permissions: API tokens grant access to all available shared models; fine-grained model-level access is unsupported.
- Stateless Endpoint: The API stores no session state or conversation context; context history must be sent with each request.
- Shared Cluster Variance: Shared tier latency fluctuates based on multi-tenant cluster demand and strict RPM/TPM rate limits.
Asset historyActive 2 of the last 12 weeksTMUpdatedNo updates · 1 bar = 1 week i
- ?Name not public?Name not publicThe Cloud Framework team knows who this is. The name is not shown on the site.