AI & LLMs

Aleph Alpha releases Kolibri, a sovereign German-English LLM

Aleph Alpha launched Kolibri, an open-weight mixture-of-experts model trained in Europe. It prioritizes data sovereignty and efficient German language processing for self-hosted deployments.

DocBento preview

Aleph Alpha released Kolibri 1 on October 3, 2026, introducing a new open-weight large language model designed specifically for German and English tasks. Built with European data sovereignty and regulatory compliance in mind, the model offers organizations a way to run advanced AI infrastructure without relying on foreign-controlled services or cloud providers.

What happened

Kolibri is a mixture-of-experts model containing 78.1 billion total parameters, though it activates only about 3.46 billion per token during inference. This architecture allows it to deliver high performance while maintaining computational efficiency comparable to much smaller models. The weights are available under the Apache 2.0 license on Hugging Face, but Aleph Alpha retains rights to the underlying training code and methods.

The model was trained from scratch using approximately 24 trillion tokens, with more than one-fifth of the dataset consisting of German text. Training occurred entirely on infrastructure located in Germany and Finland, ensuring adherence to European and German laws. Aleph Alpha describes this approach as sovereign, meaning customers gain full freedom of deployment and intellectual-property safety, with compliance treated as an inherited property of the system.

While the core training was local, the process did incorporate some external tools. English web text was rephrased using Google’s Gemma 4, German text with Mistral-NeMo, and Qwen3-32B was used for labeling data in quality filters. The team also filtered training data to reduce political bias, a step they noted was informed by their own measurements of bias in Chinese open models.

Key details

  • Architecture: 78.1 billion total parameters with 384 experts per layer; each token activates 6 experts plus one shared expert.
  • Memory requirements: Approximately 78 GB of GPU memory is required to hold the weights in 8-bit floating point format.
  • Context window: Natively supports 262,144 tokens and has been tested up to 1,048,576 tokens using sliding-window attention.
  • Languages: Optimized exclusively for German and English, with a specialized tokenizer that reduces German token count by 11.2% compared to GPT-5.
  • Reasoning: Offers four levels of reasoning effort (none, low, medium, high) and performs strong in German mathematical benchmarks.
  • License: Apache 2.0 for weights and configuration files, allowing commercial use and modification.

Background

Mixture-of-experts (MoE) models differ from traditional dense models by dividing neural network layers into multiple smaller sub-networks, known as experts. A router mechanism selects only a few of these experts to process each specific token, rather than activating the entire network for every operation. This design allows MoE models to scale up parameter counts for knowledge retention while keeping active computation low, effectively simulating a smaller model’s speed with a larger model’s capacity.

Sovereign AI refers to the ability of a nation or organization to control its own artificial intelligence infrastructure, including data storage, model training, and inference. This concept has gained traction in Europe due to regulations like the EU AI Act, which emphasize data privacy and stewardship. By hosting models locally on private servers, organizations ensure that sensitive data never leaves their physical control, mitigating risks associated with third-party cloud providers or foreign jurisdiction.

Why it matters

For teams running self-hosted software, Kolibri provides a viable alternative to US-dominated large language models that may not comply with strict European data regulations. The model’s design allows ministries, automotive suppliers, and other regulated entities to deploy advanced AI capabilities on-premises. This setup ensures that proprietary documents and customer data remain within the organization’s secure environment, addressing key concerns around data leakage and regulatory fines.

The model’s efficiency in processing German text is particularly significant for enterprises operating in DACH regions. Standard tokenizers often break German compound words into inefficient fragments, increasing computational load and reducing context window effectiveness. Kolibri’s specialized tokenizer handles these linguistic structures natively, allowing more content to fit within the same context window and reducing the cost per token for German-language applications.

However, the hardware requirements present a barrier for smaller teams. Requiring 78 GB of GPU memory means that standard consumer hardware is insufficient, necessitating data-center-grade GPUs like NVIDIA A100s or H200s. Additionally, the model’s weaker performance in multi-turn tool calling and coding tasks suggests it is best suited for document analysis and reasoning rather than autonomous agent workflows.

What you can do

  • Evaluate your current GPU infrastructure to determine if you have at least 78 GB of VRAM available for model weights.
  • Test Kolibri’s tokenizer on your specific German legal or technical documents to quantify token savings compared to your current stack.
  • Implement retrieval-augmented generation pipelines, as Kolibri is trained to admit when it lacks information rather than hallucinating answers.
  • Configure the reasoning effort parameter based on task complexity, using low effort for simple lookups and high effort for complex analysis.
  • Monitor Aleph Alpha’s vLLM plugin updates, as the model currently requires specific inference engine versions for optimal performance.
  • Consider hybrid deployments where Kolibri handles German/English document queries while other models manage coding or multi-language tasks.

More news

All news