← News

Aleph Alpha Kolibri release cover — sovereign open-weight English-German MoE model

3 Oct 2026

Aleph Alpha

Aleph Alpha open-sources Kolibri, a 78B German sovereign MoE with 1M context

On Oct. 3, 2026 — Germany’s reunification day — Aleph Alpha released Kolibri, an English–German Mixture-of-Experts Transformer with 78.1 billion total parameters (about 3.46 billion active per token), context up to 1 million tokens, and full weights on Hugging Face under Apache 2.0 for sovereign on-premise deployment.

Europe keeps talking about sovereign AI. Kolibri is a concrete open-weight drop — bilingual MoE weights you can run on your own GPUs under Apache 2.0 — aimed at German public-sector and industrial buyers who will not ship data to a U.S. cloud chat box.

On Saturday, 3 October 2026, Aleph Alpha released Kolibri. The company blog is titled “Kolibri Has Landed: A Sovereign Open-Weight Model.” The page dates the post 3 October 2026. It does not print an hour. The opening line says the release falls on the Day of German Reunification. In Germany that holiday is 3 October. The post calls Kolibri an English–German mixture-of-experts Transformer. A Transformer is the usual design for a model that reads and writes text. Mixture of experts, shortened to MoE, means the model is split into many specialist pieces, called experts. For each small piece of text, a token, only some of those experts run. A parameter is one learned number inside the model. Those lines are Aleph Alpha’s.

How big it is. The comparison table on the blog lists 78.1 billion total parameters and 3.46 billion active parameters per token. Active per token means the slice that actually does the work for each piece of text. The rest of the numbers stay in memory. The opening paragraph rounds that to 3 billion active out of 78 billion total. The Hugging Face model card, Aleph-Alpha/Kolibri-1, prints the exact counts: 78,103,074,560 total parameters and 3,457,573,120 active per token. The card says each layer has 384 experts, that 6 of them are chosen for a token, and that one more expert is always on. The blog’s table lists the same 384 total and 6 active. Those figures are Aleph Alpha’s.

How much text it can hold. The blog says Kolibri supports context lengths of up to 1 million tokens. Context length is how much text the model can keep in view at once. The model card prints the ceiling as 1,048,576 tokens. It also says the model was pre-trained on sequences of 16,384 tokens, mid-trained at 65,536, and then trained at 262,144 tokens in a final long-context stage. 262,144 is the length it was trained to. The card says quality and serving have been checked up to the full 1,048,576, and it recommends 262,144 or less when speed matters or the task is hard. Serving past 262,144 takes an extra setting that raises the limit to 1,048,576. Those lines are the blog and the card. The commands are in Sources.

Where the weights are. Weights are the learned numbers, the file a team downloads to run the model. The blog says the full weights are on Hugging Face and may be used under the Apache 2.0 license. The model card names the repository Aleph-Alpha/Kolibri-1, lists the license as Apache 2.0, and dates the release 3 October 2026. Apache 2.0 is the open-source license both pages name. Those lines are Aleph Alpha’s.

Who it is for, in the company’s words. Aleph Alpha calls Kolibri a specialized language model for sovereign, mission-critical work in regulated areas, including public administration, industrials, and aerospace. Sovereign, on this page, is two things: how the model was built, and how a customer is allowed to run it. The post says the size is meant so a customer can run it on their own machines, on premise, without sending internal data to a third-party inference service. Inference is the step where a trained model writes an answer. On premise means the computers sit with the customer. Those lines are Aleph Alpha’s description of the product. They are not a list of agencies that have switched to it.

Where it was trained, and how much German it read. Aleph Alpha says its teams built the model in Germany and trained it on infrastructure in Germany and Finland, under European and German law. The company says that work had no foreign control. The blog says German is more than a fifth of the pre-training tokens, about 21.3 percent, roughly 4.3 trillion tokens inside a 20 trillion token pre-training run. It puts English at roughly 62 percent of that mix and code at roughly 14 percent. Pre-training is the long first pass that sets the model’s numbers. The company says it built a bilingual English–German tokenizer, the tool that cuts text into tokens, so German is part of the design rather than a later translation layer. The model card’s training-data line prints a different mix for that same 20 trillion tokens: about 62.5 percent English, 23.9 percent German, and 13.6 percent code. The 21.3 percent figure is the one the blog explains. The 23.9 percent figure is the one the card prints. Both are Aleph Alpha’s. They are not the same number.

The training job, and the checkpoint that stayed inside. The blog’s table says Kolibri finished pre-training on 11 September 2026. That run used 768 NVIDIA B200 graphics chips and about 20 trillion pre-training tokens, at a length of 16,384, over 21 days. A B200 is an NVIDIA data-center chip. After pre-training, the blog describes a mid-training stage of 3.44 trillion tokens at 65,536, and a long-context stage of 200 billion tokens at 262,144. It calls the three stages together nearly 24 trillion tokens. The model card’s long-context line says 201 billion tokens. The 200 billion figure is the blog’s. An earlier internal model, Kolibri Origin, is listed at 30.6 billion total parameters and 3.27 billion active per token, with pre-training finished on 11 June 2026. The table’s release cell for Origin says there was no public release. The opening paragraph rounds Origin to 30 billion parameters and a 65,000-token context. Those lines are Aleph Alpha’s. Origin is not the model on Hugging Face.

What the company says about the law, and about refusing an answer. Aleph Alpha says Kolibri was built with the EU AI Act, the General-Purpose AI Code of Practice, and the GDPR in mind, and that copyright was a focus of that work. The EU AI Act is Europe’s artificial-intelligence statute. GDPR is the European data-protection law. The model card says Aleph Alpha is a signatory of the EU general-purpose AI Code of Practice and links the European Commission page for that code. A signature there is the company’s statement that it signed. On grounding, the blog says the model was trained with abstention examples and with a method it calls Merlin-Arthur, so it can say it does not know when the context is not enough. The blog says Kolibri holds back a wrong answer on 44 percent of items in its AA-Omniscience test, against 15 percent for Kolibri Origin. The table prints 44.0 and 14.8. Those rates are Aleph Alpha’s own tests.

What the score table shows. Aleph Alpha says the benchmarks were run on its own harnesses, and that where a model has a reasoning setting, the run used the highest one. A harness is the company’s test setup. On that table Kolibri’s overall score is 75.5 in English and 70.8 in German. In the same table Qwen3.8 27B, listed as a dense model, scores 80.2 overall in English and 79.9 in German. A dense model runs nearly all of itself for every token. Aleph Alpha’s written claim is not that Kolibri leads every row. It says that across math, coding, grounding, and long-context tasks, Kolibri matches models with up to four times its active parameter count, and it names Nemotron 3 Super as an example. That model appears in the table as Nemotron 3 Super 120B-A12B, in the group labeled 12 billion active parameters. The post also says Kolibri sits on the Pareto frontier for quality versus serving cost, in English and in German. It defines that frontier as the best trade-offs between two goals, where doing better on one means giving something up on the other. It says none of the compared models delivers more quality at the same serving cost, or the same quality at a lower cost. That comparison is Aleph Alpha’s, about the models in its chart.

How a team runs it. Aleph Alpha says Kolibri needs its aleph-alpha-inference package, which supplies a plugin for vLLM. vLLM is an open program that serves a model to other software. The blog says a team can use a container image the company publishes, or install the package, which also installs the vLLM version it supports. The model card says the weight file is about 78 gigabytes, and it lists a minimum of one B200, one B300, or two H100 chips, among other chips of that class. A user can also set how hard the model thinks. The blog’s table lists four reasoning settings: none, low, medium, and high. The install line and the serve line are in Sources.

What two same-day write-ups add. TestingCatalog, dated 3 October 2026, restates the blog’s figures: 78.1 billion total parameters, 3.46 billion active per token, contexts up to 1 million tokens, full weights on Hugging Face under Apache 2.0, and the option to run the model on the customer’s own machines. It says the benchmark numbers are the vendor’s. Startup Fortune, by Judith Murphy, stamps Oct. 3, 2026, 4:07 p.m. It does not print a time zone next to that clock. It calls Aleph Alpha the Heidelberg company and says Kolibri is built to run on infrastructure the customer controls, for government and industrial buyers. It also says Aleph Alpha has not published parameter counts. The company blog and the model card do publish them. The city name is Startup Fortune’s. The blog does not print Heidelberg.

The picture is Aleph Alpha’s release cover. A green field fills the frame, lighter through the middle and deeper along the top and bottom. A white hummingbird sits to the left of the word Kolibri, in white. Faint outlined rectangles step in from the left and right edges. The frame does not print a calendar date. It is the cover graphic. It is not a photograph of a lab, a chip, or a person.

In plain terms, Aleph Alpha on Saturday put Kolibri’s full weights on Hugging Face under Apache 2.0. The company describes an English–German mixture-of-experts model with 78.1 billion parameters, about 3.46 billion of them active for each token, and a context that can be served up to about 1 million tokens. It says the model was trained in Germany and Finland, that German was about 21.3 percent of pre-training on the blog’s count, and that a customer can run it on their own machines. Pre-training finished on 11 September 2026 on 768 B200 chips. An earlier checkpoint of about 30 billion parameters, Kolibri Origin, was not released. The score table is Aleph Alpha’s. On it, Kolibri’s overall number is not the highest in the chart.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources