Updates·October 3, 2026, 22:39

Aleph Alpha releases open Kolibri with 1 million tokens of context

AI-generated and checked against the sources listed below.

The German AI company Aleph Alpha has released Kolibri, a bilingual English and German model with open weights under Apache 2.0. It is aimed at public administration and regulated industries that want to run AI on their own hardware.

AI-generated image

Aleph Alpha has released Kolibri, a model with open weights that can be downloaded in full on Hugging Face under the Apache 2.0 license. The model is bilingual (English and German) and aimed at the public sector and regulated industries, where data should preferably not leave the premises.

Kolibri is a Mixture-of-Experts model with 78.1 billion parameters in total, of which 3.46 billion are active for each token. It supports context of up to one million tokens, and it can call tools. Companies can run it on-premises instead of sending internal data to a third-party inference service.

How the model is built

Kolibri builds on the earlier Kolibri Origin and on Aleph Alpha's automated Model Factory. It was trained on 768 B200 GPUs: first 20 trillion tokens over 21 days, then mid-training and adaptation to long context, just under 24 trillion tokens in total.

The architecture has 384 experts, six of which are active per token. Only 10 of 50 layers use full attention, while the other 40 use a sliding window of 512 tokens. That is meant to keep operating costs down. The long-context adaptation itself reached 256,000 tokens, but Aleph Alpha provides settings so the model can run with 1,048,576 tokens.

German makes up 21.3 percent of the tokens in pretraining. The model has a bilingual vocabulary of 128,000 entries and a tokenizer designed to preserve German compound words. There are four reasoning settings: none, low, medium and high. According to Aleph Alpha, the model specializes in German, math, code, long context and agent tasks. There are also industry-specific evaluations for public administration, automotive, semiconductors, industrial technology and aviation, which do not use customer data.

Figures and caveats

In the company's own tests, Kolibri scores 75.5 overall in English and 70.8 in German. It gets 96.9 on AIME 2025, 85.9 on LiveCodeBench v6 and 61.4 overall on BFCL v4. Aleph Alpha says the model sits on the Pareto frontier between quality and operating cost in both languages and that it can match models with up to four times as many active parameters on math, code, grounding, agent tasks and long context.

However, these are results the company itself ran with its own test setups and the highest available reasoning setting. They should therefore be seen as an indication, not as independent documentation.

A central point is grounding. Kolibri was trained with examples of abstaining from answering and with Aleph Alpha's Merlin-Arthur method, which is meant to teach it to hold back when evidence is lacking. On AA-Omniscience, it avoided giving a wrong answer in 44 percent of cases, compared with 14.8 percent for Kolibri Origin. It reached 0.23 on the company's own M/A grounding score.

What it means

The model was built in Germany and trained in Germany and Finland. Aleph Alpha says that control over data selection, training, evaluation, weights and deployment is designed to meet European requirements for compliance and sovereignty. Deployment uses the company's own inference package and a Kolibri-specific vLLM plugin.

For Danish companies in regulated industries, the point is that a model with open weights can be run internally and that it is optimized for German and English. Danish is not mentioned as a supported language, so that needs to be tested before you count on it. For ordinary AI users, the model is primarily interesting as a sign that European alternatives with open licenses are on the way.

What it means for you

Kolibri is first and foremost for you if you work with IT or data at a company or authority where sensitive information must not leave the premises. The model is free to download and can be run on your own hardware, so your IT department can try it out without sending data to a third party. The model is built for English and German, so test it thoroughly in Danish before you use it for anything important. If you just use AI for everyday things, it is mostly a sign that more European and open alternatives are coming.

Sources

More on this topic

Get the week's AI news in your inbox

Choose your level, topics and length. One email a week, unsubscribe at any time.

Subscribe to Promptly Newsletter
PromptlyNewsletterRSSLog in

The news on aijour is AI-generated and checked against the cited sources.