Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI announces that it now delivers up to 2x faster serverless inference for leading open-source LLMs, ranking #1 in output speed benchmarks. The performance improvements come from coordinated enhancements across next-gen GPU hardware, optimized kernels, near-lossless quantization, and production-grade speculative decoding with custom-trained draft models.
From the source
Together AI now delivers up to 2x faster serverless inference for demanding open-source LLMs, ranking #1 in output speed benchmarks. The performance breakthrough comes from coordinated improvements across next-gen GPU hardware, optimized kernels, near-lossless quantization, and production-grade speculative decoding with custom-trained draft models.
together.ai