# Hugging Face — Bamba: Inference-Efficient Hybrid Mamba2 Model

- Company: Hugging Face (huggingface.co)
- Announced: 2024-12-18T00:00:00+00:00
- Category: new-model
- Subject: Platform
- Open weights: yes
- Models affected: Bamba-9B
- Source: https://huggingface.co/blog/bamba
- Record: https://forck.live/items/1779-bamba-inference-efficient-hybrid-mamba2-model

IBM, Princeton, CMU, and UIUC announce Bamba-9B, a hybrid Mamba2 model trained on 2.2T tokens with open data, demonstrating 2.5x throughput improvement and 2x latency speedup over standard transformers in vLLM, and released with support in transformers, vLLM, TRL, and llama.cpp.

## Evidence

Verbatim from https://huggingface.co/blog/bamba:

> We introduce Bamba-9B, an inference-efficient Hybrid Mamba2 model trained by IBM, Princeton, CMU, and UIUC on completely open data.

---

Record: https://forck.live/items/1779-bamba-inference-efficient-hybrid-mamba2-model
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
