# Kyutai — CASA: The return of the cross-attention

- Company: Kyutai (kyutai.org)
- Announced: 2025-12-23
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://kyutai.org/casa
- Record: https://forck.live/items/18350-casa-the-return-of-the-cross-attention
- Subject: Moshi / Pocket TTS

Moritz Böhle*, Amélie Royer*, Juliette Marrie*, Edouard Grave, and Patrick Pérez. We revisit cross-attention (CA) as a fusion mechanism for vision-language models (VLMs), a design that has been largely put aside in favor of token insertion in recent year, despite CA’s favorable efficiency properties for streaming understanding tasks. While several early multimodal models relied on cross-attention, current SotA VLMs instead use direct token insertion, where image embeddings produced by a vision encoder are directly interleaved with text tokens in the language model stream. This allows full interaction between modalities, but it also causes memory and compute costs to grow with the number of images processed, as image tokens accumulate in the model’s KV-cache. In contrast, cross-attention injects visual information through additional dedicated CA layers, where text tokens (queries) attend to image tokens (keys and values). By design, when extending cross-attention to a text-image interleaved sequence containing multiple images, we only use the latest seen image as key-value of the cross-attention layers. …

---

Record: https://forck.live/items/18350-casa-the-return-of-the-cross-attention
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
