# TII — Welcome Falcon Mamba: The first strong attention-free 7B model

- Company: TII (tii.ae)
- Announced: 2024-08-12T12:00:00+00:00
- Category: new-model
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://falcon-lm.github.io/blog/falcon-mamba/
- Record: https://forck.live/items/18418-welcome-falcon-mamba-the-first-strong-attention-free-7b-model
- Subject: Falcon LLM
- Open weights: yes
- Models affected: Falcon Mamba
- License: TII Falcon Mamba 7B License 1.0

TII released Falcon Mamba, a 7B-parameter pure Mamba architecture model trained on ~5500GT of data. The model can process sequences of arbitrary length without increased memory storage and fits on a single A10 24GB GPU. It is open access under the TII Falcon Mamba 7B License 1.0 and available via Hugging Face.

## Evidence

Verbatim from https://falcon-lm.github.io/blog/falcon-mamba/:

> Falcon Mamba is based on the original Mamba architecture, proposed in Mamba: Linear-Time Sequence Modeling with Selective State Spaces , with the addition of extra RMS normalization layers to ensure stable training at scale. This choice of architecture ensures that Falcon Mamba: can process sequences of arbitrary length without any increase in memory storage, in particular, fitting on a single A10 24GB GPU.

---

Record: https://forck.live/items/18418-welcome-falcon-mamba-the-first-strong-attention-free-7b-model
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
