
Ai2 Olmo-core 3 Targets Trillion-Parameter MoE Training
Published by AINave Editorial
Ai2’s Olmo-core 3 targets a central tension in mixture-of-experts training: only some experts work on each token, but the full expert pool and training state still have to fit across a GPU cluster. The Allen Institute for AI says its framework can scale the expert pool from eight to 128 while selecting four experts per token, with infrastructure intended for models above one trillion parameters. That is a scaling ambition, not evidence here of a completed trillion-parameter training run. Ai2’s framework and its stated scale
The memory problem is bigger than active computation
A mixture-of-experts (MoE) model activates selected portions of its parameters for each token, unlike a dense model that uses the entire model. That can reduce computation per token, but it does not remove the need to store the full model across GPUs or coordinate expert activity during training. The distinction between active computation and the full model footprint
Olmo-core 3 addresses that footprint by dividing work and state across devices. Expert parallelism places different experts on different GPUs; splitting model layers across GPU groups reduces how much of the model each card needs to hold. A distributed optimizer also spreads optimizer state, the data used to calculate training updates, rather than keeping a full copy on every GPU. The framework’s distribution techniques
Ai2 also supports MXFP8, a lower-bit number format that the report says can reduce computation and data movement between GPUs. These techniques target different parts of the training burden: model weights, optimizer state, and the traffic involved in moving data. The value is in coordinating those pieces, rather than expecting one format change to solve the whole memory problem. MXFP8 support and its stated purpose
The throughput figure has a narrow frame
Ai2 reported 52,000 tokens per second for a 47-billion-parameter model on Nvidia B3000 GPUs. SiliconANGLE compared that with about 19,400 tokens per second for Nvidia Megatron-Core, describing the result as roughly 2.7 times the throughput. The reported comparison does not establish that the same gain holds for other model sizes, hardware, or workloads. The reported throughput figures and test context
That distinction matters because throughput is only one measure of a training system. Olmo-core 3’s architectural case is that it distributes model and optimizer state so large MoEs can be trained without each GPU holding everything. The benchmark offers a concrete result for one stated configuration; the available details do not establish how performance changes across other deployments.
The project and related systems are available to developers and the open-source community on GitHub, according to SiliconANGLE. Olmo-core 3’s reported availability The consequential question for teams is whether its distribution strategy and performance carry over to their own model and cluster configuration.






















