Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
fal presents experiments on overlapping communication and computation in the pre-attention stage of Ulysses context parallelism for video diffusion models. The post benchmarks Async Ulysses, Async Ulysses with Symmetric Memory, and Fused QKV projections on an 8xB200 GPU node, reporting chunk latency reductions of up to 37.3% and end-to-end improvements of up to 5.0% at lower GPU counts.
From the source
Async Ulysses does what we want: chunk latency drops by about 23–25% at 2/4/8 GPUs, while end-to-end improves by ~3%.
blog.fal.ai