TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Aleph Alpha has released Kolibri, a German-English mixture-of-experts model whose full weights are available under the Apache 2.0 license. The company reports a context window of up to one million tokens and strong results on selected benchmarks, but independent verification, deployment requirements and the model’s practical performance remain to be established.
Aleph Alpha has released Kolibri, an open-weight German-English language model with 78 billion total parameters, of which 3 billion are active, and a context window of up to one million tokens. The company says the full weights can be downloaded from Hugging Face under the Apache 2.0 license, making the release relevant to organizations seeking to run AI systems on their own infrastructure rather than rely solely on outside inference services.
Aleph Alpha describes Kolibri as a mixture-of-experts Transformer specialized for German, reasoning, mathematics and agentic tasks. It is intended for regulated and mission-critical settings, including public administration, industrial companies and aerospace. The company says its design aims to balance model capability with serving costs by activating 3 billion of its 78 billion parameters during use. That architecture figure does not by itself establish the hardware, memory or operating costs required for a particular deployment.
The company says Kolibri was trained through the same pipeline as its earlier Kolibri Origin, a 30-billion-total, 3-billion-active-parameter model with a 65,000-token context window. Aleph Alpha describes the process as covering data curation, pre-training, post-training and evaluation, and says it used hundreds of ablation experiments. It also reports that training could continue through hardware failures or interrupted data connections without a person stepping in. These are company descriptions of its development process, not independently verified findings in the supplied material.
Aleph Alpha’s published benchmark table reports results across mathematics, knowledge, coding, long-context and agentic tasks, including AIME 2025, GPQA Diamond, LiveCodeBench and LongBench Pro. The company says Kolibri matches models with as much as four times its active parameter count on a range of tasks. Its table shows mixed results across individual measures and competitors, so the broad performance claim should be read alongside the benchmark names, scores and evaluation conditions rather than as a universal ranking.
Local Deployment for Regulated Work
Open weights can give organizations more control over where a model runs and how it is adapted, but they do not automatically provide a complete, locally deployable system. Aleph Alpha says Kolibri is designed to run on premises, which could help customers keep internal data within their own environments. That may matter to public agencies and regulated businesses facing requirements around data handling, intellectual property and operational oversight.
The release also puts a German-developed model into a market where organizations weigh language performance, cost, compliance and supplier dependence. Aleph Alpha frames sovereignty as both the model’s development chain and the ability to transfer it to customers. Whether that translates into practical independence depends on details such as compute access, licensing in specific uses, technical support and the customer’s own infrastructure. The company’s claims about measurable return on investment are objectives, not reported results from customer deployments in the supplied announcement.
As an affiliate, we earn on qualifying purchases.
From Kolibri Origin to Release
Kolibri follows Aleph Alpha’s Kolibri Origin, which the company presents as an earlier validation of its model-training pipeline. Origin had 30 billion total parameters, 3 billion active parameters and a 65,000-token context window; the new model increases the stated context capacity to one million tokens. Aleph Alpha says its iterative pipeline shortened the time between the two releases, though the announcement does not give a detailed schedule for training or development.
The company says it created internal evaluation suites for customer-related fields including the German public sector, aviation, manufacturing and automotive. It reports proxy-benchmark scores moving from 0.72 to 0.99 for an automotive supplier, 0.35 to 0.80 for semiconductors, and 0.54 to 0.70 for the German public sector. These are Aleph Alpha’s internal evaluations, not results from named customer deployments; the supplied material does not specify enough about their methodology to compare them directly with public benchmarks.
“Kolibri is an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active.”
— Aleph Alpha
open-weight language model software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Independent Tests and Deployment Needs
The supplied announcement does not include an independent evaluation of Kolibri’s benchmark results or a full account of testing conditions, such as inference hardware, decoding settings and data contamination controls. It is also not clear from the material how much compute and memory customers will need to serve the model, what throughput to expect in production, or how the one-million-token context performs on specific workloads.
Aleph Alpha refers readers to a technical report, but its detailed methods and results are not included in the source material provided here. The company’s internal proxy scores and claims about sovereignty, compliance and return on investment therefore remain company-reported. Actual legal compliance depends on the deployment and its use, not on an open-weight license alone.
AI infrastructure for regulated industries
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Technical Report and Customer Trials
The next useful evidence will be the technical report referenced by Aleph Alpha, along with independent tests that disclose evaluation methods and deployment costs. The company says customers can download the weights from Hugging Face, but the supplied announcement does not name launch customers or provide production case studies. Public evaluations and documented deployments could clarify how Kolibri performs in German-language workflows, long-context tasks and regulated environments.
Organizations considering the model will need to check the license, technical requirements and governance of their intended use, then test performance against their own workloads. No further release date, customer rollout schedule or independent certification is specified in the source material.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Kolibri?
Kolibri is Aleph Alpha’s German-English mixture-of-experts language model, with 78 billion total parameters and 3 billion active parameters, according to the company.
Can organizations download and use the model?
Aleph Alpha says the full weights are available on Hugging Face under the Apache 2.0 license. Users should review the license and deployment requirements for their specific use.
What does the one-million-token context window mean?
The stated maximum is the amount of text the model can process in one context. The announcement does not establish how well it handles every workload at that length or what serving resources are required.
Have Kolibri’s benchmark claims been independently verified?
The source material presents Aleph Alpha’s benchmark results. It does not provide independent verification or enough methodological detail to confirm the company’s comparisons.
Who is Kolibri intended for?
Aleph Alpha says it is designed for regulated and mission-critical uses, including public administration, industrials and aerospace, with an emphasis on German-language and domain-specific work.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
