# TII — Falcon‑H1R-FP8: Accelerating Inference with Quantized Precision

- Company: TII (tii.ae)
- Announced: 2026-02-16T08:00:00+00:00
- Category: model-update
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://falcon-lm.github.io/blog/falcon-h1r-7b-fp8/
- Record: https://forck.live/items/18411-falcon-h1r-fp8-accelerating-inference-with-quantized-precision
- Subject: Falcon LLM
- Models affected: Falcon H1R 7B FP8, Falcon H1R 7B

TII released Falcon H1R 7B FP8, a fully quantized version of the Falcon H1R 7B model with both weights and activations in FP8 format. The model preserves BF16-level accuracy (AIME25 drops 0.8%, LCB-v6 falls 1%, GPQA-D shows 0.1% difference) while achieving 1.2×–1.5× throughput improvement and reducing GPU memory footprint by about 44% (from 14.2 GB to 7.9 GB). The quantization was performed using NVIDIA Model Optimizer with per-tensor FP8 post-training quantization and also quantizes the KV cache to FP8.

## Evidence

Verbatim from https://falcon-lm.github.io/blog/falcon-h1r-7b-fp8/:

> Using NVIDIA Model Optimizer and post-training quantization (PTQ) workflow, the FP8 quantized model preserves the original BF16 quality performance while delivering a 1.2×–1.5× throughput boost and halving GPU memory footprint.

---

Record: https://forck.live/items/18411-falcon-h1r-fp8-accelerating-inference-with-quantized-precision
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
