# fal — Ulysses Unbound: Experiments in Communication–Computation Overlap

- Company: fal (fal.ai)
- Announced: 2026-02-23T18:23:51+00:00
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://blog.fal.ai/ulysses-unbound-experiments-in-communication-computation-overlap/
- Record: https://forck.live/items/8004-ulysses-unbound-experiments-in-communication-computation-overlap
- Subject: fal platform

fal presents experiments on overlapping communication and computation in the pre-attention stage of Ulysses context parallelism for video diffusion models. The post benchmarks Async Ulysses, Async Ulysses with Symmetric Memory, and Fused QKV projections on an 8xB200 GPU node, reporting chunk latency reductions of up to 37.3% and end-to-end improvements of up to 5.0% at lower GPU counts.

## Evidence

Verbatim from https://blog.fal.ai/ulysses-unbound-experiments-in-communication-computation-overlap/:

> Async Ulysses does what we want: chunk latency drops by about 23–25% at 2/4/8 GPUs, while end-to-end improves by ~3%.

---

Record: https://forck.live/items/8004-ulysses-unbound-experiments-in-communication-computation-overlap
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
