Germany’s New Sovereign AI Model
Key Takeaways:
- Aleph Alpha released Kolibri 1 on October 3, 2026: a bilingual German-English model with roughly 78B total parameters and 3.46B active per token, Apache 2.0 weights on Hugging Face.
- Sparse compute does not mean light deployment: FP8 weights still need roughly 78 GB in memory.
- Aleph Alpha’s own benchmarks put Kolibri ahead on German math, but behind Qwen on tool-calling and long-context tests. Those comparisons are company-run, not independently verified.
- The release lands mid-merger, as Aleph Alpha folds into Cohere in a deal valued at roughly $20 billion.
What Aleph Alpha Actually Shipped
On October 3, 2026, German Reunification Day, Aleph Alpha released Kolibri 1, a bilingual German-English model with full weights on Hugging Face under the Apache 2.0 license. It was trained on infrastructure in Germany and Finland, and shipped alongside a 189-page technical report. The announcement came three months after its predecessor, Kolibri Origin, finished pre-training on June 11, 2026.

Kolibri is the first model Aleph Alpha has released publicly. Kolibri Origin, a 30.6B-parameter model with a 65k-token context window, stayed internal as validation of the training pipeline. In three months, the team increased the model size from 30B to roughly 78B parameters, expanded the context window from 65k tokens to a validated 1M, and grew the training tokens from 7.5 trillion to 20 trillion.
The release comes during restructuring. Cohere and Aleph Alpha signed a definitive merger agreement on September 16, 2026, after disclosing plans in April, per a joint press release. TechCrunch reported the combined company would be valued at roughly $20 billion, citing the FT, with Schwarz Group investing $600 million into Cohere’s Series E round, per CNBC. The combined business operates as Cohere, dual-headquartered in Berlin and Toronto.
Kolibri is the first major product from that restructured organization, targeting public administration, industry, and aerospace rather than consumer chatbots. The focus is on compliance-heavy, infrastructure-sensitive work for buyers who cannot route documents through a hyperscaler in Virginia.
Inside the Architecture
Kolibri is a mixture-of-experts transformer. The model card lists about 78.1B total parameters, with about 3.46B active per token. Each of its 50 layers contains 384 routed experts plus one shared expert, and the router sends every token through six of the 384. That roughly 4.4% activation ratio lets Kolibri compute like a much smaller model while keeping the capacity of a much larger one.

The attention design reduces the cost of long contexts. Only 10 of 50 layers process the full context; the other 40 use a sliding window of 512 tokens, so their decode cost stays limited regardless of input length. Rotary position embeddings apply only in the sliding-window layers, which allows the model to extend beyond its trained context without position scaling. Native context length is 262,144 tokens, with quality validated up to 1,048,576.
Pre-training ran on 768 NVIDIA B200 GPUs (96 HGX 8xB200 nodes) for 21 days over 20 trillion tokens, followed by mid-training of 3.44 trillion tokens at 64k sequence length and long-context adaptation of about 200 billion tokens at 256k, totaling roughly 24 trillion tokens, nearly three times Kolibri Origin’s. The router uses a technique Aleph Alpha calls exact quantile balancing, an extension of the quantile-balancing approach introduced in Kimi K3; the company says exact computation, which Kimi K3 considered too expensive, improves both load balance and model quality at fixed cost.
Post-training ran in two stages. Supervised fine-tuning used 174 billion tokens of synthetic data filtered into a 268-billion-token mix, and reinforcement learning covered more than 1.2 million curated tasks across code, math, agentic tool use, and instruction following. The model reasons at four effort levels (none, low, medium, high), so one deployment can trade latency against quality per request.
Serving Kolibri with vLLM
The model card specifies vLLM as the serving library with Aleph Alpha’s plugin. FP8 weights ship as float8_e4m3fn in 128×128 blocks with dynamically quantized activations, while embeddings, the language-model head, norms, and the MoE router remain in bfloat16. A minimal deployment:
# Serve Kolibri 1 with vLLM + the Aleph Alpha plugin
# Note: this assumes the plugin is installed and FP8 weights are downloaded.
# Production should pin the vLLM build, set --max-num-seqs to bound
# concurrency, and reserve KV-cache memory for long-context requests.
vllm serve Aleph-Alpha/Kolibri-1 \
--tensor-parallel-size 2 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-reasoning \
--reasoning-parser kolibri \
--served-model-name kolibri-1
The --max-model-len 262144 flag follows Aleph Alpha’s guidance: the model can reach 1,048,576 tokens, but the company recommends contexts of at most 262,144 for latency- or throughput-sensitive deployments and complex tasks. The 78 GB FP8 footprint means a single H200 or B200 holds the weights, but the minimum is 2x A100 80GB or 2x H100 SXM5, with 2x H100 SXM5 recommended for production.
German Tokenizer and German Data
German is a core focus. Roughly 21.3% of pre-training tokens are German, and only about 6% of the data came from translation. Aleph Alpha’s writing on German training data argues heavy reliance on machine translation produces a model carrying the source language’s cultural fingerprint, so the team curated German web data directly and converted existing German documents into multiple formats.
Aleph Alpha required about 4 trillion German tokens to reach its 20% target at the 20-trillion-token horizon. After deduplication and filtering, open German datasets provided only about 390 billion tokens. The team filled the gap with a German-specific Common Crawl pipeline (1.3 trillion unique tokens) and by converting existing German documents into encyclopedia entries, Q&A, and passage formats (about 1 trillion tokens). Translation, used only in Kolibri Origin, was dropped due to “translationese” artifacts.
The tokenizer adds to the data advantage. German forms long compound words that English-focused tokenizers split into pieces. An independent review by developer Tejas Kumar tested six tokenizers against Germany’s Basic Law. Kolibri needed 35,190 tokens for the German text, while OpenAI’s o200k_base tokenizer (used by GPT-4o and GPT-5) required 41,482, about 17.9% more. On the official English translation, the two were nearly equal. Kolibri’s tokenizer has a 128,000-token vocabulary trained with a byte-pair-encoding variant Aleph Alpha calls UniBPE, which uses a Unigram objective for merge selection. Fewer tokens per German text means more document fits in a fixed window and fewer decode steps per page, though Kumar treats this as a tokenizer test, not a quality test.
Benchmarks: Leads and Gaps
Aleph Alpha reports strong German math results. On AIME 2025 translated to German, Kolibri scores 87.5 versus 82.9 for Qwen3.6-35B-A3B and 85.6 for Nemotron 3 Super 120B-A12B. On English AIME 2025, Kolibri’s 96.9 beats Nemotron’s 91.7. The company credits the German math lead to training the model to reason in German, detailed in a separate post about “cold-starting” German reasoning.
The results differ on other tests. Qwen3.6-35B-A3B leads Kolibri on BFCL v4 tool calling (67.2 to 61.4) and LongBench Pro (70.8 to 64.5). On the Artificial Analysis Omniscience index, Kolibri’s -32.8 trails Qwen’s -15.3, meaning it has more factual gaps from memory, though it is trained to abstain rather than guess.
| Benchmark | Kolibri | Qwen3.6-35B-A3B | Nemotron 3 Super 120B-A12B |
|---|---|---|---|
| AIME 2025 (German) | 87.5 | 82.9 | 85.6 |
| AIME 2025 (English) | 96.9 | 84.6 | 91.7 |
| BFCL v4 (tool calling) | 61.4 | 67.2 | 61.0 |
| LongBench Pro | 64.5 | 70.8 | 62.9 |
| AA-Omniscience Index | -32.8 | -15.3 | -36.5 |
Two notes. First, these comparisons were run by Aleph Alpha, not an independent evaluator, so treat them as company-reported. Second, the “matches models with up to four times its active parameter count” claim compares active, not total, parameters. Qwen3.6-35B-A3B has 35B total parameters to Kolibri’s 78B, so where Qwen wins, it wins with less than half the stored weights.
Grounding is a clearer, separate result. Aleph Alpha trained Kolibri with its Merlin-Arthur protocol, where the model must answer from visible evidence or admit it cannot. On the Omniscience test, when Kolibri did not know an answer, it abstained or answered partially 44% of the time, versus 11.1% for Qwen3.5 35B-A3B and 23.7% for GPT-OSS 120B; only Qwen3.6 35B-A3B did better, at 56.7%. That abstention rate is the feature Aleph Alpha emphasizes for regulated buyers, and the one benchmark result that directly supports the sovereignty pitch.
Deployment Reality
Sparse activation reduces compute, not memory. All 78 billion weights must stay in memory even though only 3.46 billion are used per token. In FP8, the weights alone take up roughly 78 GB, a point the model card states clearly: the full model must be held in memory even though only part of it is active at any time.
The sovereignty claim also has limits. Apache 2.0 covers weights and config files; Aleph Alpha retains rights to its training code and methods, so this is open-weight, not fully open development. And “sovereign” does not mean nothing external was used: the model card states English web text was rephrased with Google’s Gemma 4, German text with Mistral-NeMo, and Qwen3-32B labeled data for quality filters. Buyers needing a fully closed supply chain should weigh that against the on-premise freedom they receive.
There is also a real tradeoff. A model that abstains aggressively is safer but weaker at open-ended knowledge work, and Kolibri’s Omniscience score of -32.8 reflects that trade: it knows less from memory than a same-scale Qwen model. For public-administration use anchored in retrieval over an organization’s own documents, that is the right trade. For a general-purpose assistant, it is a drawback.
On the alternative side, Mistral holds the closest comparable position in Europe, relying on EU data-residency rules and government contracts, with its models increasingly available through sovereign hybrid deployments. Open-weight Chinese models remain the cost benchmark German enterprises quietly test against. Kolibri’s narrower focus is German depth and compliance, not leaderboard position.
Infrastructure and Merger
Kolibri did not arrive alone. Fraunhofer’s Institute for Computer Graphics Research commissioned the “Baltic Brain” GPU cluster on August 11, 2026, a three-node, 24-NVIDIA B200 system co-located inside the Green IT data center, per Tech Times. That cluster provides German public research a Blackwell-class compute base that does not route through a hyperscaler, the public-sector counterpart to the 768 B200 GPUs Aleph Alpha used for pre-training.
The commercial layer is the Cohere merger. The joint press release says the combined company will advance a partnership with Schwarz Group companies to deliver sovereign AI on STACKIT, Schwarz Digits’ sovereign cloud. Schwarz Group reported 185.6 billion euros in sales for its 2025 fiscal year, and the combined entity is expected to exceed 1,000 employees across both continents.
For a German ministry or regulated manufacturer, that combination is the actual product: a model it can run on its own servers, a cloud it can host in, and a vendor structure that keeps deployment inside European legal jurisdiction. Whether that outweighs a Qwen model’s lower cost and stronger tool-calling scores is a procurement decision, not a benchmark one.
What to Watch
The main question is adoption, not capability. Aleph Alpha has pitched Kolibri to German public administration and regulated industry, and the merger with Cohere brings global reach plus the STACKIT sovereign cloud partner. Whether German agencies route real workloads through it depends on procurement cycles that move slower than model releases.
The measurable sign over the next two quarters is whether third parties publish independent evaluations of Kolibri on standard harnesses, and whether the combined Cohere entity reports named government or industrial deployments on STACKIT. Until then, the benchmark table remains a company’s own scorecard.
Related Reading
More in-depth coverage from this blog on closely related topics:
- Yeonjun Attends Miu Miu Fashion Show
- How to Create Apple Passes with Pass Designer
- Pentagon Data Breach Reveals Security Gaps
- Linux Kernel Security Issues
- Understanding SaaS and Cloud Economics
Sources and References
Sources cited while researching and writing this article:
Rafael
Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...
