# Together AI — Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets

- Company: Together AI (together.ai)
- Announced: 2026-06-02T00:00:00+00:00
- Category: api-release
- Subject: Inference platform
- Models affected: MiniMax M3, M3
- Context window: 1M-token context window
- Source: https://www.together.ai/blog/serving-minimax-m3-for-efficient-inference-unlocking-1m-token-context-and-multimodality-without-regrets
- Record: https://forck.live/items/2343-serving-minimax-m3-for-efficient-inference-unlocking-1m-token-context-and

Together AI announces it is the preferred cloud partner for MiniMax M3 and will host the open-weights model as a developer endpoint upon its public release. The company's Inference and Kernel teams delivered optimizations including a KV-Block-Major sparse attention kernel, paged attention integration for MSA, an index scoring kernel, and a Rust-based multimodal preprocessing gateway, achieving 81–125% throughput improvements across concurrency levels. The model supports a 1M-token context window and native multimodality.

## Evidence

Verbatim from https://www.together.ai/blog/serving-minimax-m3-for-efficient-inference-unlocking-1m-token-context-and-multimodality-without-regrets:

> Together AI is the preferred cloud partner for MiniMax M3. Together AI will host the open-weights model as a developer endpoint upon its public release. Our Inference and Kernel teams delivered significant engineering breakthroughs to serve M3 efficiently, including key optimizations such as a KV-Block-Major sparse attention kernel, a novel paged attention integration for MSA, highly optimized index scoring kernel and a Rust-based multimodal preprocessing gateway, resulting in 81–125% throughput improvements across different concurrency levels.

---

Record: https://forck.live/items/2343-serving-minimax-m3-for-efficient-inference-unlocking-1m-token-context-and
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
