Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The blog post tells the story of Together AI's kernels team, highlighting their work on FlashAttention, the ThunderKittens library, and their rapid optimization of kernels for NVIDIA's Blackwell GPUs, achieving up to 2x speedups over cuBLAS on H100s within a week of hardware access.
From the source
Within one week of hardware access, we had some of the fastest FP4 and FP8 GEMM kernels available for Blackwell, with up to 2x speedups over cuBLAS on H100s.
together.ai