Europe's most persistent "sovereign AI" champion is back with its biggest technical swing yet. Aleph Alpha, the Heidelberg-based startup that has spent years positioning itself as Europe's answer to the US and Chinese frontier labs, has released Kolibri — an open-weight Mixture-of-Experts language model for English and German, with weights published on Hugging Face under the Apache 2.0 license.
The architecture numbers are the story's hook: 78.1 billion total parameters with only 3.46 billion active per token, and a context window stretched to 1 million tokens. The MoE design means inference cost scales with the active parameters rather than the total, while the million-token window targets the long-document, large-codebase and extended-conversation workloads where context length is the binding constraint. The company says the model was optimized for long context and reasoning efficiency together, supports explicit reasoning modes and tool calling, and offers four selectable reasoning depths (none, low, medium, high) at inference time — a dial for trading cost against quality.
What separates Kolibri from another open MoE release is how deliberately German it is. The team built a custom bilingual tokenizer, UniBPE, designed to represent long German compound words far more efficiently than conventional tokenizers, and kept native German data in the training mix throughout, so that German accounts for 21.3% of pretraining tokens. Aleph Alpha says training ran on infrastructure in Germany and Finland, and that the model was designed from the ground up around the EU AI Act, the General Data Protection Regime's copyright constraints and the AI code of conduct — compliance as a pre-training design constraint rather than a launch-day press release.
The most unusual feature is a hallucination countermeasure called the Merlin-Arthur protocol, which trains the model to answer "I don't know" honestly when the evidence isn't there, rather than confabulate. And in a rarity for model launches of any origin, the company published a 189-page technical report detailed enough that community reviewers have described it as a tutorial on building a modern agentic LLM — dataset creation methods included.
The caveats are real, though, and mostly company-reported. In Aleph Alpha's own evaluations, dense models like Qwen3.8 27B outscore Kolibri on some tests, and critics note the official comparisons lean on models from about a year ago rather than current lightweight MoE competitors — so the benchmark picture remains unproven by independent evaluation. Practically, MoE economics cut both ways: despite the tiny active parameter count, running Kolibri requires holding all 78 billion parameters in GPU memory — roughly 78GB of VRAM — which puts it out of reach of consumer laptops pending quantized releases the community is already asking for. There are also unconfirmed reports of merger talks with Canada's Cohere, which skeptics say would undercut the European sovereignty pitch if ownership shifts to Toronto.
Still, the release matters beyond one model. Europe now has a genuinely open, fully documented, EU-regulation-aware long-context model that organizations unwilling to send data to US or Chinese APIs can inspect, self-host and audit under a permissive license. Whether Kolibri's benchmarks hold up under independent scrutiny will decide if it becomes infrastructure or a statement — but as Mistral's French-anchored success showed, the market for "capable and jurisdictionally clean" is real, and Aleph Alpha just made its strongest bid yet to own it.
Comments (0)
Log in to join the discussion
Log InNo comments yet