Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
TRL now supports co-located vLLM, allowing training and inference to share the same GPUs, reducing idle time and improving throughput without extra hardware.
From the source
This approach is what we refer to as colocation. Training and inference are co-located on the same GPUs and coordinated via the same process group, allowing them to take turns smoothly — no extra hardware needed.
huggingface.co