⏱️ Reading time: 10 min

A model that activates just 4.4% of its own parameters per token has just emerged from Germany under an open license. It’s called Aleph Alpha Kolibri, launched by the German company Aleph Alpha on October 3, 2026, and it combines 78.1 billion total parameters with a mixture-of-experts architecture trained entirely on European infrastructure.

📑 En este artículo
  1. TL;DR
  2. What Is Aleph Alpha Kolibri
  3. What Happened
  4. Context and History
  5. Technical Details
  6. Impact and Analysis
  7. What’s Next
  8. Frequently Asked Questions
    1. What does it mean for Aleph Alpha Kolibri to be a sovereign model?
    2. How many parameters does Kolibri activate per token?
    3. What license does the sovereign German LLM have?
    4. Does Kolibri’s mixture-of-experts system need less memory than an equivalent dense model?
    5. Why is Kolibri’s tokenizer relevant for German speakers?
  9. References

The launch comes with a 189-page technical report and an Apache 2.0 license for the weights, right as Europe debates how to compete in artificial intelligence without depending on US or Chinese infrastructure.

TL;DR

  • Aleph Alpha launched Kolibri on October 3, 2026, a German LLM with 78.1 billion parameters.
  • The model activates only 3.46 billion parameters per token (4.4%) thanks to its mixture of experts.
  • It was trained on 24 trillion tokens using 768 NVIDIA B200 GPUs, located in Germany and Finland.
  • Its UniBPE tokenizer cuts the tokens needed for German by 11.2% compared to GPT-5’s tokenizer.
  • The company signed the EU’s Code of Practice for General-Purpose AI Models.

What Is Aleph Alpha Kolibri

Aleph Alpha Kolibri is an open-weight language model in German and English, with a mixture-of-experts architecture of 78.1 billion total parameters that activates only 3.46 billion per token, released under the Apache 2.0 license and designed to comply with the European AI Act from the ground up.

The name isn’t accidental: Kolibri means hummingbird in German, a nod to a model whose main merit is weighing little in compute even though it carries a considerable amount of memory. According to the model card published on Hugging Face, the model supports reasoning at four levels (none, low, medium, and high), tool calling, and has a knowledge cutoff of June 18, 2026.

What Happened

Aleph Alpha, headquartered in Heidelberg, presented Kolibri on October 3, 2026, alongside a launch post, a 189-page technical datasheet, and the weights on Hugging Face. In that post, the company describes the model with a word it repeats several times: “sovereign.” In its own words, quoted in the technical analysis published on tej.as, “teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control.”

That sovereignty has two sides, according to the same post. The first is the process: German teams, infrastructure in Germany and Finland, European law. The second is what the customer gets: full deployment freedom and intellectual property security, so regulatory compliance is built in by design and doesn’t depend on an outside provider that could change or shut down the model. Aleph Alpha also signed the European Union’s Code of Practice for General-Purpose AI Models, the voluntary framework that accompanies the AI Act for general-purpose models.

In Aleph Alpha’s internal evaluation, Kolibri outperforms all comparable models of its size in German and English. That figure should be read with caution: it comes from the maker itself, not from an independent benchmark.

Context and History

Aleph Alpha was founded in 2019 as one of Europe’s bets on building its own language models, in an ecosystem dominated by US labs and, more recently, Chinese ones. The company had previously launched the Luminous family, and Kolibri arrives as an architectural leap: from dense models to mixture of experts, with the focus explicitly placed on European regulatory compliance by design rather than as a later patch.

The underlying debate isn’t new. An independent analyst who follows the European AI scene documented in his analysis that, at a 2024 conference in the United States, he heard the phrase “you regulate, you don’t innovate” when he mentioned he lived in Germany. Kolibri works as a direct response to that criticism: an attempt to show that regulation and technical capability aren’t mutually exclusive.

Sovereign doesn’t mean isolated. Kolibri’s model card acknowledges that the English text used for training was rewritten with Google’s Gemma 4, the German text with Mistral-NeMo, and that Qwen3-32B labeled data for quality filters. Aleph Alpha also filtered out political bias that, it claims, it measured in open Chinese models. Kolibri’s sovereignty lies in the training infrastructure, the applicable law, and control over the final model, not in every byte of data having originated in Europe.

Technical Details

The core mechanism of the sovereign German LLM is the mixture of experts. Kolibri has 50 layers, each with 384 experts plus 1 shared expert that every token passes through. A router decides, for each token, which of the 384 experts to send it to: it picks 6. That combination of 6 selected experts plus the shared expert is what turns 78.1 billion total parameters into just 3.46 billion of actual compute per token.

Kolibri was trained on 768 NVIDIA B200 GPUs in data centers in Germany and Finland. Foto de Numan Ali en Unsplash

The trick has a cost that the model card itself doesn’t hide: “the full model must be held in memory even though only part of it is active at any time.” Even though the model computes like one with 3.46 billion parameters, it needs the memory of one with 78.1 billion: around 78 GB in 8-bit floating point (FP8). That’s why running Kolibri locally still requires serious hardware, even when the compute cost per token is low.

flowchart TD
 A["Input token"] --> B["Router"]
 B --> C["Shared expert"]
 B --> D["6 of 384 selected experts"]
 C --> E["Combined output"]
 D --> E

The second mechanism is the tokenizer. German joins words into long compounds, and a tokenizer trained mostly on English cuts them into not-very-useful pieces. Aleph Alpha trained its own tokenizer, UniBPE, with a vocabulary of 128,000 tokens: it combines the bottom-up merging of byte-pair encoding (BPE) with a different scoring rule, the Unigram objective, which better respects how words are built in German.

o200k_base (GPT-5): Bund | es | ver | fass | ungs | gericht -> 6 tokens
Kolibri: Bundes | verfassungsgericht -> 2 tokens

That word is the German name for the Federal Constitutional Court. The GPT-4o and GPT-5 tokenizer (o200k_base, via OpenAI’s tiktoken) splits it into 6 tokens; Kolibri’s leaves it at 2. According to the technical report, cited in the tej.as analysis, UniBPE needs on average 11.2% fewer tokens to process German text than GPT-5’s tokenizer, the best result among the 9 tokenizers they compared.

Kolibri has a native context of 262,144 tokens, tested up to 1,048,576. To confirm that figure without relying on what the inference runtime lets you force, it’s worth checking directly the config.json file in the Hugging Face repository and the field that sets the maximum position length: that number is the real limit of the architecture, not the one negotiated by the server that exposes it. Regarding the four reasoning levels, on the other hand, there’s no direct way to verify which one is active from outside: Aleph Alpha had not published, as of this article, its own public endpoint for Kolibri, so the parameter that controls the level depends on whatever inference engine is used to serve it.

Reasoning levelWhen to use itAdvantageLimitation
NoneDirect answers, simple tasksMinimal latencyNo explicit reasoning
LowOne- or two-step logic questionsBalance of cost and speedCan fail on multi-step problems
MediumTasks that require breaking down a problemBetter accuracy on complex tasksMore output tokens, higher cost
HighLong reasoning, math, planningMaximum accuracyHigher latency and compute consumption

Impact and Analysis

The European open-weight system solves a concrete problem for governments and regulated companies: a health record, an automotive part design, or an administrative file can go in and out of a company’s own server, under German law, without an outside provider being able to change the model out from under the customer. For a ministry or a European automotive supplier, that matters more than a benchmark percentage point.

⚠️ Heads up: The comparison that puts Kolibri ahead of other models of its size is an internal evaluation by Aleph Alpha, published in its own technical report. There are still no independent benchmarks confirming those results.

The memory cost, however, puts a real ceiling on that promise of easy sovereignty. A model that needs 78 GB of memory to activate only 3.46 billion parameters per token doesn’t run on a laptop or a modest server: it still requires data-center-class GPUs, even though the compute cost per query is low. Data sovereignty doesn’t equal affordable hardware sovereignty.

There’s also tension between the rhetoric and the actual data practice. The model card is transparent in admitting it used Google’s Gemma 4 and Mistral-NeMo to rewrite training text, and Qwen3-32B to label filtering data. A model presented as a guarantee of independence from non-European infrastructure used, in its own curation process, models from Google, Mistral, and a Chinese lab. That doesn’t invalidate the technical achievement, but it does qualify how clean the line really is between “sovereign” and “with outside help.”

What’s Next

What’s missing now is external verification. No independent benchmarking lab had published, as this article went to press, an evaluation of Kolibri that checks the figures Aleph Alpha reports. With the weights already on Hugging Face under Apache 2.0, the community can run its own tests, fine-tune the model with its own data, and above all, put the central promise to the test: that a model trained and served under European law can compete at the frontier without giving up control.

The other front to watch is institutional adoption. Signing the EU’s Code of Practice is a gesture, not a purchase contract; what will matter in the coming months is how many European ministries or industrial suppliers replace foreign models with Kolibri in production, and under what infrastructure conditions.

Kolibri’s router sends each token to only 6 of its 384 experts. Foto de Hitesh Choudhary en Unsplash

Try it yourself: Kolibri’s weights and technical datasheet are already published on Hugging Face under the Apache 2.0 license, so you can review the 189-page report before deciding whether to run it on your own infrastructure.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What does it mean for Aleph Alpha Kolibri to be a sovereign model?

It means it was trained on German and Finnish infrastructure under European law, and that a customer can deploy it on their own servers without an outside provider controlling or being able to disable the model.

How many parameters does Kolibri activate per token?

It activates 3.46 billion of the 78.1 billion total, or 4.4%, because its router sends each token to only 6 of 384 experts plus a shared expert.

What license does the sovereign German LLM have?

The weights and configuration files were released under the Apache 2.0 license. Aleph Alpha keeps the code and training methods private.

Does Kolibri’s mixture-of-experts system need less memory than an equivalent dense model?

No: even though it computes like a model with 3.46 billion parameters, it needs to keep all 78.1 billion in memory, around 78 GB in FP8, because the router can send any token to any expert.

Why is Kolibri’s tokenizer relevant for German speakers?

Because German forms long compound words, and Kolibri’s UniBPE tokenizer needs 11.2% fewer tokens than GPT-5’s to represent the same German text, which makes processing that language cheaper.

References

  • tej.as: detailed technical analysis of Kolibri’s architecture, tokenizer, and report.
  • Hugging Face: repository where Aleph Alpha published Kolibri’s weights under Apache 2.0.
  • European Commission, Digital Strategy: official framework for the AI Act and the Code of Practice for General-Purpose AI Models.
  • Wikipedia: general explanation of the mixture-of-experts (MoE) architecture.
  • Aleph Alpha: official site of the German company that developed Kolibri.

📱 Do you like this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Kevin Ache en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.