Aleph Alpha Kolibri-1: 78.1B Parameters, 3.46B Active
marktechpost.com

Aleph Alpha Kolibri-1: 78.1B Parameters, 3.46B Active

Tech News
3 min read

Published by AINave Editorial

TL;DRAleph Alpha’s Kolibri-1 is an open-weight English-German model with 78.1B total parameters but 3.46B active per token. Its million-token maximum is configurable, while the reported FP8 checkpoint still requires high-end GPU hardware.

Aleph Alpha’s Kolibri-1 makes a useful distinction between how much model it contains and how much it activates for each token. The English-German mixture-of-experts model has 78.1 billion total parameters, but activates 3.46 billion, or 4.4%, per token. It is released under Apache 2.0, with weights available on Hugging Face, and is aimed at uses such as public administration, industry and aerospace.Kolibri-1’s specifications and stated target sectors

That sparse activation may help explain its compute profile, but it does not make the model a small checkpoint. The reported FP8 weights are about 78GB. MarktechPost says they can run on one B200, B300 or H200 GPU, or two H100 SXM5 GPUs, using vLLM and dedicated Kolibri reasoning and tool-call parsers. These are reported deployment details, not an independently verified compatibility test.The checkpoint size and stated GPU requirements

A million-token maximum is not the default

Kolibri’s advertised context ceiling is 1,048,576 tokens, but the article gives 262,144 tokens as the default serving context. To request the full window, operators must set the maximum model length to 1,048,576 and override the position-embedding setting. The distinction matters: the headline maximum describes a configurable serving setup, not what a default launch automatically provides.Kolibri’s context limits and serving configuration

The architecture combines 40 sliding-window attention layers, each looking back 512 tokens, with 10 full-attention layers. The sliding-window layers use a fixed-size key-value cache, while the full-attention layers grow with context length. Aleph Alpha reports that this hybrid supports sequences four times longer than a full-attention model at matched compute. That comparison is the company’s claim, rather than an independently established result.The attention design and reported matched-compute comparison

Sparse compute does not settle model choice

The model’s selective activation is a meaningful architectural feature, but teams still need to account for the full checkpoint and the hardware needed to host it. Open weights and Apache 2.0 licensing offer deployment flexibility; they do not by themselves establish regulatory compliance or make hosting inexpensive.

Aleph Alpha’s reported evaluations also show why a single score would be a poor summary. Kolibri scored 63.4 on the English agentic average, while its reported BFCL v4 score was 61.4. Those figures measure different evaluations, and the supplied reporting does not independently verify the test methodology. For teams assessing Kolibri, the clearer takeaway is the specific trade: relatively few active parameters per token alongside a large set of weights to load, and a long context that requires explicit configuration.Reported benchmark scores and deployment details

The practical question is therefore not simply whether Kolibri can process a million tokens. It is whether a workload benefits from that configured context enough to justify the serving setup and hardware the release describes.

FAQs

Kolibri-1 is an open-weight English-German mixture-of-experts language model with 78.1 billion total parameters and 3.46 billion active per token. The release is under Apache 2.0, and its weights are available on Hugging Face.

Sources

Latest Tech News