# Together AI — Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving

- Company: Together AI (together.ai)
- Announced: 2026-03-04T00:00:00+00:00
- Category: infrastructure-release
- Subject: Inference platform
- Source: https://www.together.ai/blog/cache-aware-disaggregated-inference
- Record: https://forck.live/items/2375-cache-aware-prefill-decode-disaggregation-cpd-for-up-to-40-faster-long-context

Together AI announces cache-aware prefill–decode disaggregation (CPD), a serving architecture that separates cold and warm workloads by cache hit rate to improve throughput and reduce time-to-first-token for long-context inference.

## Evidence

Verbatim from https://www.together.ai/blog/cache-aware-disaggregated-inference:

> At Together AI, we built cache-aware prefill–decode disaggregation (CPD), a serving architecture that purposely separates cold and warm workloads by cache hit rate, resulting in fast context reuse.

---

Record: https://forck.live/items/2375-cache-aware-prefill-decode-disaggregation-cpd-for-up-to-40-faster-long-context
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
