# Together AI — How a global fintech scaled coding agent traffic with Dedicated Model Inference

- Company: Together AI (together.ai)
- Announced: 2026-09-18
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://www.together.ai/blog/global-fintech-scales-coding-agent-traffic-with-dedicated-model-inference
- Record: https://forck.live/items/12222-how-a-global-fintech-scaled-coding-agent-traffic-with-dedicated-model-inference
- Subject: Inference platform
- Models affected: GLM 5.2, GLM 5.1
- Context window: 256K context

Together AI published a case study documenting how a global fintech scaled its coding assistant workload using Dedicated Model Inference. The customer runs GLM 5.2 through Together's DMI service to handle spiky, engineering-hours traffic, and gained self-service endpoint provisioning, programmatic access to performance metrics, and the ability to swap models and adjust configuration without redeployment.

## Evidence

Verbatim from https://www.together.ai/blog/global-fintech-scales-coding-agent-traffic-with-dedicated-model-inference:

> With DMI, the customer's engineers scale endpoints, roll out models, and test changes themselves, no tickets, no waiting on Together. The result: infrastructure that moves as fast as the teams adopting it.

---

Record: https://forck.live/items/12222-how-a-global-fintech-scaled-coding-agent-traffic-with-dedicated-model-inference
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
