# Perceptron — Composing Weight and Data Sparsity in MoE

- Company: Perceptron (perceptron.inc)
- Announced: 2026-01-21
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.perceptron.inc/blog/composing-weight-and-data-sparsity-in-moe
- Record: https://forck.live/items/18618-composing-weight-and-data-sparsity-in-moe
- Subject: Perceptron Isaac / Perceptron Mk1

Improving compute efficiency through varying compute per token January 21, 2026 Composing Weight and Data Sparsity in MoE Improving compute efficiency through varying compute per token January 21, 2026 Composing Weight and Data Sparsity in MoE Improving compute efficiency through varying compute per token Weight Sparsity vs Data Sparsity MoE layers achieve efficiency through weight sparsity : each token activates only k of N experts. But this is only half the story. Think of routing as a matrix R where columns are experts and rows are tokens. Weight sparsity constrains the rows – each token uses at most k experts. Data sparsity constrains the columns – each expert processes a bounded number of tokens. These are orthogonal. Weight sparsity asks "which experts for this token?" Data sparsity asks "which tokens for this expert?" Composing both gives you a budget over the full matrix. Why does this matter? Not all tokens need equal compute. Blank image patches, punctuation, predictable continuations – these don't deserve the same budget as dense, information-rich tokens. Data sparsity lets the model skip compute where it's not needed. …

---

Record: https://forck.live/items/18618-composing-weight-and-data-sparsity-in-moe
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
