# Cartesia — Mamba‑3B-SlimPJ: State-space models rivaling the best Transformer architecture

- Company: Cartesia (cartesia.ai)
- Announced: 2023-12-14
- Category: open-weight-release
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://www.cartesia.ai/blog/mamba-3b-slimpj
- Record: https://forck.live/items/8033-mamba-3b-slimpj-state-space-models-rivaling-the-best-transformer-architecture
- Subject: Sonic / Ink
- Open weights: yes
- Models affected: Mamba-3B-SlimPJ
- License: Apache 2.0
- Context window: context length 2048

Cartesia and Together released Mamba-3B-SlimPJ, a 2.8B parameter state-space model trained on 600B tokens of the SlimPajama dataset, under an Apache 2.0 license. The model matches the performance of the strongest comparable 3B Transformer models with 17% fewer training FLOPs.

## Evidence

Verbatim from https://www.cartesia.ai/blog/mamba-3b-slimpj:

> We’re releasing the strongest Mamba language model yet, Mamba-3B-SlimPJ, in partnership with Cartesia & Together under an Apache 2.0 license. Trained on 600B tokens, Mamba-3B-SlimPJ matches the quality of some of the best 3B Transformers such as BTLM-3B-8K (also trained for 600B tokens) with 17% fewer FLOPs.

---

Record: https://forck.live/items/8033-mamba-3b-slimpj-state-space-models-rivaling-the-best-transformer-architecture
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
