# Together AI — Autoscaling endpoints for LLM inference

- Company: Together AI (together.ai)
- Announced: 2026-07-31T00:00:00+00:00
- Category: capability-change
- Subject: Inference platform
- Source: https://www.together.ai/blog/autoscaling-endpoints-for-llm-inference
- Record: https://forck.live/items/2326-autoscaling-endpoints-for-llm-inference

Together AI announces autoscaling endpoints for LLM inference on its Dedicated Model Inference platform, allowing users to configure scaling based on metrics like in-flight requests, TTFT, GPU utilization, and token throughput, with adjustable replica bounds and timing windows.

## Evidence

Verbatim from https://www.together.ai/blog/autoscaling-endpoints-for-llm-inference:

> With Dedicated Model Inference on the Together AI platform you can get your deployments to autoscale on metrics the inference engine actually understands, such as in-flight requests, TTFT, GPU utilization, token throughput.

---

Record: https://forck.live/items/2326-autoscaling-endpoints-for-llm-inference
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
