Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI announces autoscaling endpoints for LLM inference on its Dedicated Model Inference platform, allowing users to configure scaling based on metrics like in-flight requests, TTFT, GPU utilization, and token throughput, with adjustable replica bounds and timing windows.
From the source
With Dedicated Model Inference on the Together AI platform you can get your deployments to autoscale on metrics the inference engine actually understands, such as in-flight requests, TTFT, GPU utilization, token throughput.
together.ai