Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI announces new reliability and operational control features for Together GPU Clusters, including passive health checks, auto node repair, and a rebuilt Slurm-on-Kubernetes stack (Slinky 1.0) to improve failure detection, automated recovery, and cluster management for production workloads.
From the source
We've spent the last several weeks shipping a set of changes to Together GPU Clusters aimed at the operational reality of running training and inference at scale: hardware fails, schedulers leak, and teams outgrow the single-admin-kubeconfig workflow they started with.
together.ai