Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI announces it is the preferred cloud partner for MiniMax M3 and will host the open-weights model as a developer endpoint upon its public release. The company's Inference and Kernel teams delivered optimizations including a KV-Block-Major sparse attention kernel, paged attention integration for MSA, an index scoring kernel, and a Rust-based multimodal preprocessing gateway, achieving 81–125% throughput improvements across concurrency levels. The model supports a 1M-token context window and native multimodality.
From the source
Together AI is the preferred cloud partner for MiniMax M3. Together AI will host the open-weights model as a developer endpoint upon its public release. Our Inference and Kernel teams delivered significant engineering breakthroughs to serve M3 efficiently, including key optimizations such as a KV-Block-Major sparse attention kernel, a novel paged attention integration for MSA, highly optimized index scoring kernel and a Rust-based multimodal preprocessing gateway, resulting in 81–125% throughput improvements across different concurrency levels.
together.ai