Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Tencent Hunyuan research introduces the Effective Learning Rate (ELR) as a better heuristic for controlling angular update size (AUS) during LLM pretraining. The post shows that ELR leads to Hyperball optimization, enables more accurate loss prediction, and the proposed minus-square-root schedule sets a new SOTA on the Modded-nanoGPT Track 3 benchmark, reaching target loss 3.28 in 3175 steps.
From the source
When training Large Language Models (LLMs), the learning rate (LR) is often treated as the main knob controlling optimization speed. ... This motivates the notion of *angular update size (AUS)* ... In this post, we focus on directly controlling the AUS during LLM training. Specifically, we rewrite the update rule ... and explicitly control the AUS via the *effective learning rate (ELR)*
hunyuan.tencent.com