Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
Cartesia released a technical report introducing MOHAWK, a multi-stage architecture distillation method that converts Transformer models into efficient Mamba-2 variants. The company released model weights for the Llamba family (Llamba-1B, Llamba-3B, Llamba-8B), distilled from corresponding Llama-3.X models, and reported throughput improvements up to 12× for Llamba-8B versus Llama-3.1-8B. The work targets on-device and high-throughput inference scenarios.
From the source
Our latest technical report , “Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing,” describes new ideas we’re exploring in architecture distillation—a method that transforms a pre-trained model into a new, more efficient model architecture.
cartesia.ai