Research

Olmo-core 3 scales MoE training to trillions of parameters

The newly released Olmo-core 3 framework introduces a redesigned mixture-of-experts training system, allowing researchers to scale open-source AI models to trillions of parameters.

Hugging Face Blog15 hrs agoResearch
Image: Hugging Face Blog

The developers behind the Olmo language models have launched Olmo-core 3, an upgraded open-source training framework featuring a redesigned mixture-of-experts (MoE) system. To overcome the high communication and memory costs of scaling MoEs, the new framework transitions from fully sharded data parallelism (FSDP) to a distributed data parallelism (DDP) architecture. This shift keeps experts resident on the GPUs and routes data directly to them, eliminating the need to repeatedly gather and reshard model weights.

In preliminary tests on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE model achieved a throughput of 52,000 tokens per second per GPU, representing a 2.7-fold increase over the previous FSDP-based implementation's 19,400 tokens. In another benchmark, the framework scaled its expert pool from 8 to 128 while keeping active parameters per token stable at 3.2 billion. This expanded total parameter capacity from 4.6 billion to 47 billion with less than a 5% drop in training throughput. The system also successfully benchmarked a 1.2-trillion-parameter model with 58.36 billion active parameters per token across 512 GPUs, reaching a peak throughput of 858 TFLOP/s per GPU, and reached 2.38 trillion total parameters in a short-capacity test using DeepEP v2.

To achieve these speeds, Olmo-core 3 combines expert and pipeline parallelism with a distributed optimizer, rowwise expert parallelism, and GPU-resident routing. It also introduces support for the lower-precision MXFP8 format. In a controlled test on four NVIDIA B300 GPUs, enabling MXFP8 boosted training throughput by 21% compared to a BF16 baseline, while reducing peak active memory from 103 GiB to 95 GiB.

The release also highlights critical training insights for practitioners. The developers warned of "token gerrymandering," where routing balance scores improve even as actual workloads become less balanced. They also noted that lowering expert learning rates did not improve performance, and overlapping communication with computation sometimes slowed execution. This fully open-source stack will serve as the foundation for the next generation of Olmo models.

This is our own summary of reporting by Hugging Face Blog

More in Research