Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI presents FlashAttention-4, an algorithm and kernel co-design for Blackwell GPUs that maximizes overlap between matmul and resource bottlenecks like exponential units and shared memory traffic, achieving up to 1605 TFLOPs/s on B200 with BF16.
From the source
On B200 with BF16, it reaches up to 1605 TFLOPs/s (71% utilization), up to 1.3× faster than cuDNN version 9.13 and 2.7× faster than Triton.
together.ai