# Thinking Machines Lab — Modular Manifolds

- Company: Thinking Machines Lab (thinkingmachines.ai)
- Announced: 2025-09-26
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://thinkingmachines.ai/blog/modular-manifolds/
- Record: https://forck.live/items/18570-modular-manifolds
- Subject: Thinking Machines / Inkling

When we train large neural networks, we need to keep them healthy. We do not want the tensors in the network—either the weights, activations or gradients—to grow too large or too small. Very small and very large tensors cause a variety of problems not just limited to numerical underflow and overflow. For example, weight matrices changing size during training makes it harder to design training algorithms—since the relative size of updates to weights has a significant impact on the speed of learning. The gold standard for keeping tensors healthy is to normalize them. Normalization is commonplace for activation vectors, where we use techniques like layer norm to put the activations on a good scale before passing them to the next layer. It is also commonplace to normalize gradient updates, where we can interpret fast training algorithms like the Muon optimizer as spectrally normalizing the updates. Normalization provides us with certainty about the sizes of tensors—without needing to check Wandb!—and when training large neural networks with many interacting components, having certainty about the network internals is valuable. …

---

Record: https://forck.live/items/18570-modular-manifolds
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
