THE SIGNAL IN ONE SENTENCE
An open model release changes who gets to ask the hard questions. Before the weights are downloadable, the developer decides which prompts to run, which failures to show, which comparison models to include and how much access anyone else receives. After release, an independent lab can put the same machine on a different track. Aleph Alpha opened that door on October 3 with Kolibri, an English and German mixture-of-experts reasoning model released with downloadable weights under Apache 2.0 terms. The model card lists 78,103,074,560 total parameters, with 3,457,573,120 active for each token. That distinction is the main engineering trick. Kolibri carries the capacity of a much larger model in memory but routes each token through only a small selection of its expert components. The result is meant to reduce the computation required for each answer without pretending the rest of the model disappears. The memory bill still arrives. Aleph Alpha says the FP8 release occupies about 78 gigabytes and requires at least two 80-gigabyte A100s, two H100 SXM5 units, one H200, one B200 or one B300. A research university, federal computing center or large company may be able to operate that stack. A small German municipality is not slipping it under a receptionist's desk next to the label printer. That is the first useful distinction in this release. Open weights mean an organization can download the learned parameters and run the model on infrastructure it controls. Open weights do not mean the system is cheap, fully reproducible, trained from public data or independently proven. Kolibri is more open than a hosted chat box. It is less complete than a scientific reproduction kit. Aleph Alpha has published unusually specific model documentation. The company says Kolibri was trained from scratch on 20 trillion pretraining tokens, followed by 3.44 trillion tokens in mid-training and 201 billion tokens for long-context extension. It describes the pretraining mixture as roughly 62.5 percent English, 23.9 percent German and 13.6 percent code. It reports 6.4 times 10 to the 23 floating-point operations and identifies the principal training hardware as 768 Nvidia B200 accelerators. Those numbers give researchers something to inspect. They do not identify the underlying corpus item by item, establish permission for every source or let another team recreate the data pipeline. The published license boundary deserves the same careful reading. The Hugging Face model card says the released weights and configuration files use Apache 2.0. It also says that the rights apply only to those published artifacts and do not extend automatically to other code, training methods or intellectual property not included in the repository. A buyer should therefore inventory each component it plans to use rather than treating one familiar license label as a blanket permission slip for the entire development stack. That is not a reason to dismiss the release. It is a reason to describe it accurately. Kolibri is designed around German and English rather than trying to be equally mediocre in a parade of languages. Aleph Alpha says its tokenizer uses 4.7 bytes per token for German and 4.2 for English. The model supports an explicit reasoning mode, tool calling, retrieval-augmented work, coding and structured extraction. Its native long-context training reached 262,144 tokens. Aleph Alpha reports validating quality and serving efficiency out to 1,048,576 tokens, while recommending no more than 262,144 for complex or performance-sensitive use. The million-token headline is real as a stated tested limit. It is not a promise that every fact buried in a million-token file will be found, interpreted correctly and cited reliably. Long context describes how much material the system can accept. Retrieval quality describes whether it finds the right part. Grounding describes whether the answer stays attached to that evidence. Latency and memory determine whether anybody can afford to repeat the job. Those measurements belong together. The company's benchmark tables are ambitious. Aleph Alpha reports strong results across mathematics, science questions, coding, tool use, long context, grounding and German-language tasks. It says Kolibri scored 81.3 percent on the German version of GPQA Diamond, 90 percent on the German AIME 2026 set and 94.7 percent on the telecom portion of Tau2-Bench. Those are developer-reported results. The release says the benchmarks were run with Aleph Alpha's own harnesses and, where applicable, the highest available reasoning effort for each model. That may be a reasonable testing choice. It also means the company selected and operated the comparison environment. The contextualized evaluations are even closer to home. Aleph Alpha built application suites for areas including the German public sector, semiconductors, aerospace, automotive suppliers and industrial drive technology. The company says these suites were informed by customer conversations about failures, workflows and difficult cases. It says customer data was not used for training and that the environments were randomized to improve robustness. This is sensible product engineering. Customers report where a system fails, the developer turns those failures into evaluations and future versions are measured against them. It is also a home-field advantage. The developer knows the test design, the scoring choices, the prompting harness and the kinds of failure that shaped the suite. Outside model teams may receive only the published score. A private evaluation can be valuable for customers while remaining weak evidence for a broad public claim. The answer is not to throw out every company benchmark. The answer is to label its evidentiary weight. A useful benchmark record should identify the exact model snapshot, precision, hardware, inference software, prompts, tool definitions, sampling settings, number of runs, scoring code, exclusions and confidence intervals. If the test cannot be released because it contains customer-derived material, an independent evaluator can run it under confidentiality and publish the method, aggregate result and conflict statement. Without that layer, the table tells us what Aleph Alpha observed in Aleph Alpha's laboratory. It does not yet tell us what a German ministry, law firm, manufacturer, hospital or university will observe after installing the model in its own environment. That local test is where Kolibri becomes interesting. German organizations have practical reasons to want a bilingual model they can operate privately. Public agencies work with administrative language, records, forms and rules that do not always travel cleanly through an English-first system. Manufacturers hold technical manuals and incident records that cannot be scattered across an ordinary consumer service. Legal and research teams need citations, access controls and stable model versions. Regulated organizations may need to choose where data, logs and operators live. A downloadable model can support those controls. It cannot provide them alone. The institution still needs a secure inference service, identity rules, document permissions, logging, retention limits, a vulnerability process, model updates, evaluation gates and a named human who owns the decision. If Kolibri is connected to tools, the surrounding system needs least-privilege credentials and a review step before consequential action. Aleph Alpha's own model card recommends human review rather than unreviewed autonomous action. That sentence should travel with every deployment diagram. The best first use is not a universal government brain. It is a bounded workflow with a measurable baseline. A ministry could test whether the model extracts fields from one family of forms while preserving links to the source page. A manufacturer could compare technical-manual answers against a retrieval system and require citations for every claim. A legal team could evaluate clause extraction on a hidden set reviewed by lawyers. A research library could measure German and English question answering separately, including abstention when evidence is missing. Each pilot should compare Kolibri with at least one smaller local model, one hosted alternative and the current human workflow. Measure accuracy, unsupported claims, latency, energy, hardware utilization, review time and cost per completed task. Then publish the failures. The model card already acknowledges systematic bias, outdated knowledge, political bias, errors and the possibility that users may mistake generated material for human work. Those are category labels, not local evidence. A deployment team needs examples from its own documents, dialects, policy rules and users. German strength should also be tested outside translated benchmark questions. Administrative compounds, legal references, regional language, mixed German and English, abbreviations, scanned correspondence and incomplete sentences are ordinary operating conditions. A model can score well on a formal test and still make a mess of the email that begins with three unexplained acronyms and ends with a photo of a fax. Independent testing should begin with five questions. Can another team reproduce the public benchmark scores from the released model and documented settings? Does Kolibri outperform a smaller model when both use the same retrieval system and hardware budget? Does the million-token mode improve real tasks enough to justify its latency and memory cost? Does the model know when the supplied evidence is insufficient, especially in German? Can an organization update or replace the model without rebuilding the entire application? Those questions turn sovereignty from a slogan into operations. Aleph Alpha calls Kolibri a sovereign open-weight model. The downloadable weights improve technical control because customers are not required to send every request to a vendor-hosted endpoint. The bilingual focus gives German institutions a locally relevant option. The public model card provides more detail than many commercial releases. But sovereignty is not created by the nationality of the developer or the location of a server. It depends on who controls the weights, data, keys, logs, administrators, evaluation, update schedule and exit path. It also depends on whether the institution has the people and hardware to exercise that control rather than merely owning a large file it cannot safely operate. Kolibri gives Europe a serious artifact to test. Now the most valuable work moves outside Aleph Alpha's laboratory. Download the weights. Rebuild the public evaluations. Test German tasks that were not used to shape the model. Compare the full operating cost. Publish the cases where the model abstains, invents, slows down or needs a human rescue. The release opened the engine bay. Independent evidence has to take the car off the home track.
01
WHAT ACTUALLY CHANGED
Aleph Alpha released Kolibri on October 3 with downloadable FP8 weights and configuration files under Apache 2.0 terms.
The model card reports 78.1 billion total parameters, 3.46 billion active parameters per token, German and English support, reasoning mode, tool calling and a validated context limit of 1,048,576 tokens.
Aleph Alpha disclosed training-mixture percentages, major training phases, compute, hardware requirements and detailed developer-run benchmark results.
The company also published contextualized application evaluations informed by customer failure cases, while stating that customer data was not used for training.
02
WHY THIS MATTERS
Downloadable weights let European institutions inspect, host and preserve a capable bilingual model without routing every request through a vendor service.
The model still requires substantial accelerator memory, deployment software, security controls and skilled operators, so possession is not the same as practical independence.
Most published performance evidence comes from the model maker using its own harnesses, and several sector evaluations are not available for independent reproduction.
Kolibri offers a concrete test of whether European sovereign AI can be measured through operational control, portability and reproducible evidence rather than geography or branding.
03
WHERE IT COULD HELP
- Run German and English document extraction on controlled infrastructure while preserving links to the source record.
- Build retrieval systems for technical manuals, public rules, contracts or research collections with task-specific citation and abstention tests.
- Evaluate tool-calling workflows inside a restricted sandbox with least-privilege credentials and required human approval before consequential action.
- Compare Kolibri with smaller local and hosted models on one hidden institutional test set, including cost, latency, energy, unsupported claims and reviewer time.
- Create an exit drill that proves weights, prompts, retrieval data, evaluations and logs can move to another operator or model without rebuilding the service.
KEEP A HAND ON THE WHEEL
Aleph Alpha and its model card establish the October 3 release, architecture, parameter counts, German and English focus, context limits, training summary, hardware requirements, Apache 2.0 weight license and developer-run benchmark results. The model card says the licence applies to the published weights and configuration files, not automatically to every unlisted artifact or training method. The company says its contextualized suites reflect customer failure cases, use randomized environments and do not train on customer data. The reviewed sources do not provide an independent benchmark reproduction, a complete item-level training corpus, public versions of all customer-derived application suites, measured production outcomes, full deployment cost, energy per task or evidence that every organization can operate the required hardware. Watch for outside reproductions, public German task suites, transparent harness configurations, energy and cost measurements, adversarial tests, deployment incident reports and proof that institutions can switch models without losing their applications or records.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
Mixture of experts
A model architecture with many specialist components that activates only a selected few for each input.
OPEN GLOSSARY CARD
Context window
The amount of input and generated material a model can consider during one call.
OPEN GLOSSARY CARD
Quantization
Representing model numbers with lower precision to reduce memory use and sometimes improve inference efficiency.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on October 5, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US