Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI announced major improvements to its Batch Inference API, including a streamlined UI, support for all serverless models and private deployments, a 3000x rate limit increase (from 10M to 30B tokens), and 50% cost reduction compared to the real-time API.
From the source
Rate limits are up from 10M to 30B enqueued tokens per model per user, a 3000× increase.
together.ai