From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Upgrades MoE training to trillion-parameter scale
Hugging Face released Olmo-core 3, an upgrade to its framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system.
The system is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency.
In one benchmark, increasing the expert pool from 8 to 128 while keeping active parameters per token fixed at about 3.2B resulted in total parameter capacity growing from 4.6B to 47B with training throughput falling by less than 5%.
From the source
Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency.
huggingface.co