Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI introduces consistency diffusion language models (CDLM), a method that accelerates diffusion language model inference by combining consistency-based multi-token finalization with block-wise KV caching, achieving up to 14.5x latency speedups on math and coding tasks.
From the source
We introduce consistency diffusion language models (CDLM), which accelerates diffusion language model inference by combining consistency-based multi-token finalization with block-wise KV caching, achieving up to 14.5x latency speedups on math and coding tasks.
together.ai