Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
GaLore is a method that reduces memory footprint for training large language models by projecting gradients into low-rank subspaces, enabling training of up to 7 billion parameter models on consumer GPUs like the NVIDIA RTX 4090. It achieves over 82.5% reduction in memory for optimizer states and can be combined with 8-bit optimizers for further savings.
From the source
“more than 82.5% reduction in memory for storing optimizer states during training”
huggingface.co