From the source
Ai2 released Olmo-core 3, an upgrade to its open-source framework for training large language models, featuring a redesigned mixture-of-experts (MoE) training system.
The framework is designed to scale MoE training into the trillion-parameter range and includes techniques such as expert parallelism, pipeline parallelism, and a distributed optimizer.
In benchmarks, a 47-billion-parameter MoE achieved 2.7× the throughput of the previous implementation, and MXFP8 support improved training throughput by about 21% over BF16.






