Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI introduces AutoJudge, a research method that accelerates LLM inference through task-specific lossy speculative decoding, which automatically identifies which generated tokens affect downstream quality without manual annotation. It achieves 1.5–2x speedups over standard speculative decoding by accepting up to 40 draft tokens per verification cycle with slight accuracy drop, and will be presented at NeurIPS 2025.
From the source
We introduce AutoJudge, a method that accelerates large language model (LLM) inference through task-specific lossy speculative decoding.
together.ai