Ai2 Olmo-core 3 Targets Trillion-Parameter MoE Training
siliconangle.com

Ai2 Olmo-core 3 Targets Trillion-Parameter MoE Training

Tech News
3 min read

Published by AINave Editorial

TL;DRAi2’s Olmo-core 3 is an open training framework designed to reduce the GPU memory burden of large mixture-of-experts models. Its reported 2.7× throughput comparison is tied to a 47-billion-parameter model and Nvidia B3000 GPUs, not a universal performance claim.

Ai2’s Olmo-core 3 targets a central tension in mixture-of-experts training: only some experts work on each token, but the full expert pool and training state still have to fit across a GPU cluster. The Allen Institute for AI says its framework can scale the expert pool from eight to 128 while selecting four experts per token, with infrastructure intended for models above one trillion parameters. That is a scaling ambition, not evidence here of a completed trillion-parameter training run. Ai2’s framework and its stated scale

The memory problem is bigger than active computation

A mixture-of-experts (MoE) model activates selected portions of its parameters for each token, unlike a dense model that uses the entire model. That can reduce computation per token, but it does not remove the need to store the full model across GPUs or coordinate expert activity during training. The distinction between active computation and the full model footprint

Olmo-core 3 addresses that footprint by dividing work and state across devices. Expert parallelism places different experts on different GPUs; splitting model layers across GPU groups reduces how much of the model each card needs to hold. A distributed optimizer also spreads optimizer state, the data used to calculate training updates, rather than keeping a full copy on every GPU. The framework’s distribution techniques

Ai2 also supports MXFP8, a lower-bit number format that the report says can reduce computation and data movement between GPUs. These techniques target different parts of the training burden: model weights, optimizer state, and the traffic involved in moving data. The value is in coordinating those pieces, rather than expecting one format change to solve the whole memory problem. MXFP8 support and its stated purpose

The throughput figure has a narrow frame

Ai2 reported 52,000 tokens per second for a 47-billion-parameter model on Nvidia B3000 GPUs. SiliconANGLE compared that with about 19,400 tokens per second for Nvidia Megatron-Core, describing the result as roughly 2.7 times the throughput. The reported comparison does not establish that the same gain holds for other model sizes, hardware, or workloads. The reported throughput figures and test context

That distinction matters because throughput is only one measure of a training system. Olmo-core 3’s architectural case is that it distributes model and optimizer state so large MoEs can be trained without each GPU holding everything. The benchmark offers a concrete result for one stated configuration; the available details do not establish how performance changes across other deployments.

The project and related systems are available to developers and the open-source community on GitHub, according to SiliconANGLE. Olmo-core 3’s reported availability The consequential question for teams is whether its distribution strategy and performance carry over to their own model and cluster configuration.

FAQs

It is the Allen Institute for AI’s development framework for large language models, with a redesigned training system for mixture-of-experts models. Ai2 says its infrastructure is intended to scale MoE training beyond one trillion parameters. The framework and its stated scale

Sources

Latest Tech News