Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI announces cache-aware prefill–decode disaggregation (CPD), a serving architecture that separates cold and warm workloads by cache hit rate to improve throughput and reduce time-to-first-token for long-context inference.
From the source
At Together AI, we built cache-aware prefill–decode disaggregation (CPD), a serving architecture that purposely separates cold and warm workloads by cache hit rate, resulting in fast context reuse.
together.ai