A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes
Hugging Face integrates LLM.int8() 8-bit matrix multiplication into its transformers library to reduce memory footprint of large language models without degrading performance.
