Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Sakana AI, in collaboration with NVIDIA, announces a paper accepted at ICML 2026 introducing open-source GPU kernels and a new sparse packing format (TwELL) to speed up inference and training of sparse transformer language models, achieving >20% speedups and reduced memory usage.
From the source
Excited to share our new #ICML2026 paper in collaboration with NVIDIA: "Sparser, Faster, Lighter Transformer Language Models". This work introduces new open-source GPU kernels and data formats for faster inference and training of sparse transformer language models:
sakana.ai