Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face introduces binary and scalar embedding quantization techniques to improve retrieval speed, reduce memory usage, and lower costs. The methods involve converting float32 embeddings to 1-bit (binary) or int8 (scalar) values, enabling 32x reduction in memory. A demo with 41 million Wikipedia texts demonstrates practical benefits. Experiments show up to ~96% performance retention with rescoring.
From the source
We introduce the concept of embedding quantization and showcase their impact on retrieval speed, memory usage, disk space, and cost.
huggingface.co