# Hugging Face — Universal Assisted Generation: Faster Decoding with Any Assistant Model

- Company: Hugging Face (huggingface.co)
- Announced: 2024-10-29T00:00:00+00:00
- Category: capability-change
- Subject: Platform
- Models affected: gemma-2-9b, Mixtral-8x22B-Instruct-v0.1, CodeLlama-13b, vicuna-68m, tiny_starcoder_py, Qwen2-0.5B-Instruct, Llama-3.1-70B, Phi-3-medium-128k-instruct
- Source: https://huggingface.co/blog/universal_assisted_generation
- Record: https://forck.live/items/1802-universal-assisted-generation-faster-decoding-with-any-assistant-model

Hugging Face and Intel Labs introduce Universal Assisted Generation (UAG), a method that extends assisted generation (speculative decoding) to work with any pair of target and assistant models, regardless of tokenizer compatibility, enabling 1.5x-2.0x speedup for LLM inference.

## Evidence

Verbatim from https://huggingface.co/blog/universal_assisted_generation:

> Universal Assisted Generation: Faster Decoding with Any Assistant Model

---

Record: https://forck.live/items/1802-universal-assisted-generation-faster-decoding-with-any-assistant-model
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
