# Cartesia — Llamba: scaling distilled recurrent models for efficient language processing

- Company: Cartesia (cartesia.ai)
- Announced: 2025-03-05
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.cartesia.ai/blog/llamba-distillation
- Record: https://forck.live/items/8023-llamba-scaling-distilled-recurrent-models-for-efficient-language-processing
- Subject: Sonic / Ink
- Open weights: yes
- Models affected: Llamba-1B, Llamba-3B, Llamba-8B

Cartesia released a technical report introducing MOHAWK, a multi-stage architecture distillation method that converts Transformer models into efficient Mamba-2 variants. The company released model weights for the Llamba family (Llamba-1B, Llamba-3B, Llamba-8B), distilled from corresponding Llama-3.X models, and reported throughput improvements up to 12× for Llamba-8B versus Llama-3.1-8B. The work targets on-device and high-throughput inference scenarios.

## Evidence

Verbatim from https://www.cartesia.ai/blog/llamba-distillation:

> Our latest technical report , “Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing,” describes new ideas we’re exploring in architecture distillation—a method that transforms a pre-trained model into a new, more efficient model architecture.

---

Record: https://forck.live/items/8023-llamba-scaling-distilled-recurrent-models-for-efficient-language-processing
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
