# Together AI — Together AI delivers fastest inference for the top open-source models

- Company: Together AI (together.ai)
- Announced: 2025-12-01T00:00:00+00:00
- Category: capability-change
- Subject: Inference platform
- Models affected: GPT-OSS-20B, GPT-OSS-120B, Qwen-3-235B-Instruct, Qwen-3-Coder-480B, Kimi-K2-Instruct, DeepSeek-R1, DeepSeek-V3.1
- Source: https://www.together.ai/blog/fastest-inference-for-the-top-open-source-models
- Record: https://forck.live/items/2399-together-ai-delivers-fastest-inference-for-the-top-open-source-models

Together AI announces that it now delivers up to 2x faster serverless inference for leading open-source LLMs, ranking #1 in output speed benchmarks. The performance improvements come from coordinated enhancements across next-gen GPU hardware, optimized kernels, near-lossless quantization, and production-grade speculative decoding with custom-trained draft models.

## Evidence

Verbatim from https://www.together.ai/blog/fastest-inference-for-the-top-open-source-models:

> Together AI now delivers up to 2x faster serverless inference for demanding open-source LLMs, ranking #1 in output speed benchmarks. The performance breakthrough comes from coordinated improvements across next-gen GPU hardware, optimized kernels, near-lossless quantization, and production-grade speculative decoding with custom-trained draft models.

---

Record: https://forck.live/items/2399-together-ai-delivers-fastest-inference-for-the-top-open-source-models
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
