Aleph Alpha released Kolibri 1 on 3 October 2026, German Unity Day: a mixture-of-experts language model whose full weights can be downloaded from Hugging Face under the Apache 2.0 license. The model has 78.1 billion parameters, of which 3.46 billion are activated for each token, and is declared for German and English only. The Heidelberg-based company accompanies it with a technical blog post, a 189-page technical report and a public summary of the training data. As of the morning of 6 October, the Hugging Face model card listed 2,453 downloads over a rolling thirty-day period and 631 likes.
The release comes seventeen days after the signing, on 16 September, of a definitive combination agreement with Canada's Cohere. For a European public or industrial buyer, the practical question is therefore what the license guarantees, regardless of which shareholder will steer the next versions.
3.46 billion active parameters, 78 GB in memory
According to the model card, Kolibri stacks 50 layers, each with 384 routed experts and one shared expert; for each token, six routed experts are selected in addition to the shared expert. The technical report puts the share of parameters engaged at 4.4%. Only 10 layers read the full context; the other 40 are limited to a 512-token window, which caps compute and memory as the text grows longer. The native context length is 262,144 tokens. Aleph Alpha says it has validated quality up to 1,048,576 tokens but recommends not exceeding 262,144 tokens for latency- or throughput-sensitive deployments and for complex tasks.
The trade-off is spelled out in the model card itself: the full model must be held in memory, even though only a small fraction of it is active. The FP8 weights take up about 78 GB. The announced minimum configuration ranges from two 80 GB A100s or two H100 SXM5s to a single H200, B200 or B300; the recommended configuration is two H100 SXM5s, two H200s, one B200 or one B300. The blog post justifies the chosen size by serving cost. An experimental 123-billion-parameter configuration, tested during development, handled only three simultaneous 256k-token requests on two H100s, against 18 for the released version, which also decodes 28% faster.
Deployment relies on an in-house vLLM plugin, aleph-alpha-inference, published on GitHub under the same Apache 2.0 license.
A German and English corpus, without French
Pre-training covered 20 trillion tokens: about 62.5% English, 23.9% German and 13.6% code, according to the model card. The blog post, for its part, gives 21.3% German, or about 4.3 trillion tokens seen after upsampling. On top of this come 3.44 trillion tokens in the mid-training phase and 201 billion for context extension. According to the model card, the tokenizer trained on this mix compresses German at 4.7 bytes per token.
Training used 768 Nvidia B200 GPUs: 21 days and 392,000 GPU hours for pre-training alone. Energy consumption for these phases, excluding post-training, is estimated at 950 MWh. The model's implicit knowledge stops at 18 June 2026.
The model card declares only two languages and presents this scope as "a deliberate choice of depth over breadth". For a French administration or company, this scope determines the choice of model. Mistral Small 4, by comparison, claims dozens of languages, including French, and accepts images as input. Any deployment of Kolibri on French documents therefore requires prior testing.
| Kolibri 1 | Mistral Small 4 | Qwen3.6 35B-A3B | |
|---|---|---|---|
| Total parameters | 78.1B | 119B | 35B |
| Active parameters per token | 3.46B | 6.5B | 3B |
| Native context | 262,144 tokens | 262,144 tokens (256k) | 262,144 tokens |
| Inputs | text | text and image | text and image |
| Declared languages | German, English | dozens, including French | not detailed in the model card |
| License | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| Overall score English / German (Aleph Alpha measurement) | 75.5 / 70.8 | 63.1 / 61.4 | 71.4 / 67.3 |
Sources: Hugging Face model cards for Kolibri 1, Mistral Small 4 and Qwen3.6 35B-A3B; scores taken from Aleph Alpha's evaluation table.
Scores measured by the vendor, on its own benchmark
Aleph Alpha compares Kolibri with about a dozen open models, including Qwen3.6 35B-A3B, Nemotron 3 Super 120B-A12B, Mistral Small 4 119B-A6B and GPT-OSS 120B. The blog post states that the benchmarks "were run using our own harnesses", with each model set, where applicable, to its highest reasoning effort. On the overall average, Kolibri scores 75.5 in English and 70.8 in German, ahead of Nemotron 3 Super (73.0 and 67.9) and GPT-OSS 120B (72.3 and 70.2).
The same table shows unfavorable gaps. On SWE-Bench Verified, Kolibri scores 66.4 against 73.8 for Qwen3.6 35B-A3B; on BFCL v4, 61.4 against 67.2. In factual error correction (RGB Fact-Check), it drops to 34.0, while Nemotron 3 Super reaches 90.0. The dense Qwen3.8 27B model, grayed out in the table, scores 80.2 in English and 79.9 in German. The five sector-specific suites highlighted are internal to the vendor. On the one dedicated to the German public sector, Kolibri scores 75.0, against 80.0 for Qwen3.5 35B-A3B and 78.0 for Nemotron 3 Super.
Apache 2.0: irrevocable rights on the weights only
The Apache License 2.0 grants each user a "perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable" right to reproduce, modify, sublicense and distribute the work (Section 2). It adds a patent license that terminates only if the licensee itself sues over the work for patent infringement (Section 3). Redistribution requires keeping the license and attribution notices (Section 4). Trademarks are not granted (Section 6) and the work is provided "AS IS", without warranty (Section 7).
For a buyer, a copy of the weights downloaded today therefore remains usable, modifiable and redistributable whoever owns Aleph Alpha in the future. The license, however, guarantees neither subsequent versions, nor maintenance, nor support. The "License and terms" section of the model card also narrows the scope. The rights apply only to the weights and configuration files in the repository. The underlying code, model architecture, parameter settings and training methods are excluded: Aleph Alpha retains all rights to them.
The "Responsible Use" paragraph, which refers to the practices prohibited by Article 5 of the AI Act, is worded as a request addressed to users. Aleph Alpha states that it is a signatory of the European code of practice for general-purpose AI models and publishes a training data summary following the Commission's template. Its 5 October press release adds that all training data was screened against a blocklist of more than 4.5 million URLs, drawing on sources including the European Commission's Piracy Watch List.
In June, the dependency on a model served via API became concrete. On 12 June, the US government subjected Fable 5 and Mythos 5 to export controls requiring access to be restricted for foreign nationals. Unable to verify nationality in real time, Anthropic suspended both models for all its users, according to the account published by the company. The controls were lifted on 30 June: Fable 5 became available again to all users on 1 July, while Mythos 5 was initially restored only for a set of US organizations. Weights held locally are not exposed to this kind of cut-off.
"No foreign control" and a pending combination with Cohere
The 3 October blog post states that the teams built the model in Germany and trained it on infrastructure in Germany and Finland, "under European and German law, with no foreign control". It does not mention Cohere. According to the table in the blog post, Kolibri's pre-training ended on 11 September, five days before the agreement.
Cohere announced on 16 September the signing of a definitive combination agreement. The combined entity will operate under the Cohere name, with dual headquarters in Berlin and Toronto, the Heidelberg office retained as a research center, and more than 1,000 employees. Ilhan Scheer, sole CEO of Aleph Alpha since the departure of co-CEO Reto Spörri on 28 September, is to become chief operating officer of the combined company. The announcement specifies that the combined company will operate within the legal frameworks of both countries. Aleph Alpha's 5 October press release recalls the agreement and states that, until closing, the company "continues to operate independently".
The sovereignty claim describes the conditions under which Kolibri 1 was produced. If the combination goes through, subsequent versions will come from a German-Canadian company.
Kolibri, Mistral Small 4 or a Chinese model: language as the criterion
For an IT department that mainly processes German or English documents and has one B200 or two H100s, Kolibri offers a European open-weight model under a permissive license. Its compute per token is close to that of Qwen3.6 35B-A3B. For French-language corpora, Mistral Small 4 declares the language and vision support, at the cost of 119 billion parameters to host; it is the vendor the French State contracted with for its ministries.
Qwen3.6 35B-A3B is also released under Apache 2.0. The origin of this model family has nevertheless already weighed on a public-sector choice. On 23 June, according to an investigation by Le Monde, the French Treasury Directorate General (Direction générale du Trésor) shut down an internal assistant built on a Qwen model after answers deemed biased. Kolibri's model card acknowledges, for its part, that its training data contains content generated by Chinese models, and Aleph Alpha says it has reduced the political biases through filtering and dedicated alignment training. Portugal has taken a different route with Amália, a public model funded with 7 million euros.
Closing of the combination, subject to regulatory approvals, is expected before the end of the year, according to Cohere.
