# Together AI — FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling

- Company: Together AI (together.ai)
- Announced: 2026-03-05T00:00:00+00:00
- Category: research-paper
- Subject: Inference platform
- Source: https://www.together.ai/blog/flashattention-4
- Record: https://forck.live/items/2374-flashattention-4-algorithm-and-kernel-pipelining-co-design-for-asymmetric

Together AI presents FlashAttention-4, an algorithm and kernel co-design for Blackwell GPUs that maximizes overlap between matmul and resource bottlenecks like exponential units and shared memory traffic, achieving up to 1605 TFLOPs/s on B200 with BF16.

## Evidence

Verbatim from https://www.together.ai/blog/flashattention-4:

> On B200 with BF16, it reaches up to 1605 TFLOPs/s (71% utilization), up to 1.3× faster than cuDNN version 9.13 and 2.7× faster than Triton.

---

Record: https://forck.live/items/2374-flashattention-4-algorithm-and-kernel-pipelining-co-design-for-asymmetric
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
