Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The blog post reports that Infini-Attention, a method for extending context length, fails to perform well; its performance degrades with increased memory compression, and existing methods like ring attention, YaRN, and rope scaling remain superior.
From the source
Infini-attention's performance gets worse as we increase the number of times we compress the memory, and to the best of our knowledge, ring attention, YaRN and rope scaling are still the best ways for extending a pretrained model to longer context length.
huggingface.co