Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI introduces ThunderAgent, a system for high throughput agentic inference that achieves up to 2.5x higher single-node throughput and 2.4x speedup on 8 nodes by introducing a novel program abstraction for scheduling agentic workflows to mitigate KV cache thrashing.
From the source
By introducing a novel program abstraction for agentic LLM request scheduling, ThunderAgent achieves up to 2.5× higher single-node throughput in our synthetic data generation pipeline, and delivers 2.4× speedup on an 8-node cluster with near-linear throughput scaling with respect to GPU nodes.
together.ai