Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI introduces Provisioned Throughput, a reserved inference capacity for open models with token-based pricing and a 99% uptime SLA, available for MiniMax M3 and GLM-5.2 with a one-month minimum term and discounts at higher commitment levels. Costs are up to 90% below Claude Opus 4.8 at list price. Provisioned Throughput Units (PTUs) are priced at $0.05 per PTU per minute, with different burn rates for input, cached input, and output tokens. The service is available in North America, EMEA, and beyond.
From the source
We're excited to introduce Provisioned Throughput, reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA.
together.ai