From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
AsyncGRPOTrainer now syncs only LoRA adapters to vLLM across separate machines without NCCL.
Hugging Face released LoRA support in TRL's AsyncGRPOTrainer (v1.14), enabling training of LoRA adapters that sync only the adapter weights to vLLM inference servers rather than full model weights.
The implementation allows trainer and inference to run on separate Hugging Face Jobs using Storage Buckets for adapter synchronization and a proxy server for routing and broadcasting, reducing training time from 3 hours 27 minutes to 53 minutes for 500 steps in a real-world project.
From the source
LoRA support recently landed in TRL's AsyncGRPOTrainer with PR #7017, and ships with TRL v1.14. The asynchronous trainer can now train an adapter instead of the full model, and it syncs only the LoRA adapter to vLLM.
huggingface.co