# Alibaba — FlashQLA: CP-/Bwd-Friendly Fused Linear Attention Kernels for GDN

- Company: Alibaba (alibaba.com)
- Announced: 2026-04-28T02:00:00+00:00
- Category: developer-tool-release
- Subject: Qwen
- Models affected: Qwen3-Next, Qwen3-Next-80B-A3B, Qwen3.5, Qwen3.6
- Source: https://qwen.ai/blog?id=flashqla
- Record: https://forck.live/items/4606-flashqla-cp-bwd-friendly-fused-linear-attention-kernels-for-gdn

Alibaba's Qwen team open-sourced FlashQLA, a high-performance linear attention kernel library for GDN (Gated Delta Network) that achieves 2-3x forward and 2x backward speedup over the FLA Triton kernel across multiple scenarios on NVIDIA Hopper.

## Evidence

Verbatim from https://qwen.ai/blog?id=flashqla:

> Today we officially open-source FlashQLA : a high-performance linear attention kernel library built on TileLang . FlashQLA applies reasonable operator fusion and performance optimization to the forward and backward passes of GDN Chunked Prefill, achieving 2-3× forward speedup and 2× backward speedup over the FLA Triton kernel across multiple scenarios on NVIDIA Hopper.

---

Record: https://forck.live/items/4606-flashqla-cp-bwd-friendly-fused-linear-attention-kernels-for-gdn
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
