From the source
TII released Falcon H1R 7B FP8, a fully quantized version of the Falcon H1R 7B model with both weights and activations in FP8 format.
The model preserves BF16-level accuracy (AIME25 drops 0.8%, LCB-v6 falls 1%, GPQA-D shows 0.1% difference) while achieving 1.2×–1.5× throughput improvement and reducing GPU memory footprint by about 44% (from 14.2 GB to 7.9 GB).
The quantization was performed using NVIDIA Model Optimizer with per-tensor FP8 post-training quantization and also quantizes the KV cache to FP8.





