From the source
Today’s leading language models contain upwards of a trillion parameters, pretrained on tens of trillions of tokens.
Base model performance keeps improving with scale, as these trillions are necessary for learning and representing all the patterns in written-down human knowledge.
In contrast, post-training involves smaller datasets and generally focuses on narrower domains of knowledge and ranges of behavior.
It seems wasteful to use a terabit of weights to represent updates from a gigabit or megabit of training data.
This intuition has motivated parameter efficient fine-tuning (PEFT), which adjusts a large network by updating a much smaller set of parameters.
The leading PEFT method is low-rank adaptation, or LoRA.
LoRA replaces each weight matrix W from the original model with a modified version W ′ = W + γ B A W' = W + \gamma BA W ′ = W + γ B A , where B and A are matrices that together have far fewer parameters than W, and γ \gamma γ is a constant scaling factor.
In effect, LoRA creates a low-dimensional representation of the updates imparted by fine-tuning.
…





