
Aleph Alpha Kolibri-1: 78.1B Parameters, 3.46B Active
Published by AINave Editorial
Aleph Alpha’s Kolibri-1 makes a useful distinction between how much model it contains and how much it activates for each token. The English-German mixture-of-experts model has 78.1 billion total parameters, but activates 3.46 billion, or 4.4%, per token. It is released under Apache 2.0, with weights available on Hugging Face, and is aimed at uses such as public administration, industry and aerospace.Kolibri-1’s specifications and stated target sectors
That sparse activation may help explain its compute profile, but it does not make the model a small checkpoint. The reported FP8 weights are about 78GB. MarktechPost says they can run on one B200, B300 or H200 GPU, or two H100 SXM5 GPUs, using vLLM and dedicated Kolibri reasoning and tool-call parsers. These are reported deployment details, not an independently verified compatibility test.The checkpoint size and stated GPU requirements
A million-token maximum is not the default
Kolibri’s advertised context ceiling is 1,048,576 tokens, but the article gives 262,144 tokens as the default serving context. To request the full window, operators must set the maximum model length to 1,048,576 and override the position-embedding setting. The distinction matters: the headline maximum describes a configurable serving setup, not what a default launch automatically provides.Kolibri’s context limits and serving configuration
The architecture combines 40 sliding-window attention layers, each looking back 512 tokens, with 10 full-attention layers. The sliding-window layers use a fixed-size key-value cache, while the full-attention layers grow with context length. Aleph Alpha reports that this hybrid supports sequences four times longer than a full-attention model at matched compute. That comparison is the company’s claim, rather than an independently established result.The attention design and reported matched-compute comparison
Sparse compute does not settle model choice
The model’s selective activation is a meaningful architectural feature, but teams still need to account for the full checkpoint and the hardware needed to host it. Open weights and Apache 2.0 licensing offer deployment flexibility; they do not by themselves establish regulatory compliance or make hosting inexpensive.
Aleph Alpha’s reported evaluations also show why a single score would be a poor summary. Kolibri scored 63.4 on the English agentic average, while its reported BFCL v4 score was 61.4. Those figures measure different evaluations, and the supplied reporting does not independently verify the test methodology. For teams assessing Kolibri, the clearer takeaway is the specific trade: relatively few active parameters per token alongside a large set of weights to load, and a long context that requires explicit configuration.Reported benchmark scores and deployment details
The practical question is therefore not simply whether Kolibri can process a million tokens. It is whether a workload benefits from that configured context enough to justify the serving setup and hardware the release describes.






















