Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face and Intel Labs introduce Universal Assisted Generation (UAG), a method that extends assisted generation (speculative decoding) to work with any pair of target and assistant models, regardless of tokenizer compatibility, enabling 1.5x-2.0x speedup for LLM inference.
From the source
Universal Assisted Generation: Faster Decoding with Any Assistant Model
huggingface.co