Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Alibaba's Qwen team open-sourced FlashQLA, a high-performance linear attention kernel library for GDN (Gated Delta Network) that achieves 2-3x forward and 2x backward speedup over the FLA Triton kernel across multiple scenarios on NVIDIA Hopper.
From the source
Today we officially open-source FlashQLA : a high-performance linear attention kernel library built on TileLang . FlashQLA applies reasonable operator fusion and performance optimization to the forward and backward passes of GDN Chunked Prefill, achieving 2-3× forward speedup and 2× backward speedup over the FLA Triton kernel across multiple scenarios on NVIDIA Hopper.
qwen.ai