# Together AI — Benchmarking inference at scale: coding agents

- Company: Together AI (together.ai)
- Announced: 2026-05-19T00:00:00+00:00
- Category: research-paper
- Subject: Inference platform
- Models affected: Kimi K2.5
- Source: https://www.together.ai/blog/coding-agent-benchmarks
- Record: https://forck.live/items/2345-benchmarking-inference-at-scale-coding-agents

The article presents a benchmark comparing Together Inference Engine against TensorRT-LLM and SGLang on a coding agent workload using Kimi K2.5. It reports that Together Inference Engine achieves 31% higher tokens per second (TPS) and 2× better time-to-first-token (TTFT) at saturation, attributed to full-stack optimizations including ThunderMLA and custom kernel rewrites.

## Evidence

Verbatim from https://www.together.ai/blog/coding-agent-benchmarks:

> On a production coding agent workload, Together Inference Engine delivers 31% more TPS than the next fastest OSS engine on the same hardware, and maintains 2× better TTFT at saturation.

---

Record: https://forck.live/items/2345-benchmarking-inference-at-scale-coding-agents
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
