Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
Cartesia and Together released Mamba-3B-SlimPJ, a 2.8B parameter state-space model trained on 600B tokens of the SlimPajama dataset, under an Apache 2.0 license. The model matches the performance of the strongest comparable 3B Transformer models with 17% fewer training FLOPs.
From the source
We’re releasing the strongest Mamba language model yet, Mamba-3B-SlimPJ, in partnership with Cartesia & Together under an Apache 2.0 license. Trained on 600B tokens, Mamba-3B-SlimPJ matches the quality of some of the best 3B Transformers such as BTLM-3B-8K (also trained for 600B tokens) with 17% fewer FLOPs.
cartesia.ai