Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The article presents a benchmark comparing Together Inference Engine against TensorRT-LLM and SGLang on a coding agent workload using Kimi K2.5. It reports that Together Inference Engine achieves 31% higher tokens per second (TPS) and 2× better time-to-first-token (TTFT) at saturation, attributed to full-stack optimizations including ThunderMLA and custom kernel rewrites.
From the source
On a production coding agent workload, Together Inference Engine delivers 31% more TPS than the next fastest OSS engine on the same hardware, and maintains 2× better TTFT at saturation.
together.ai