From the source
Engineering July 8, 2026 · 19 min read By Emrick Sinitambirivoutin, Victor Goubet, Avshalom Manevich At H Company, our agents are served by a fleet of vLLM replicas running on GPU nodes in Kubernetes.
Agentic workloads are bursty, so being able to dynamically adjust our number of replicas is key, and that adjustment is only as fast as the time it takes a new replica to become ready.
For our largest models, this process previously took up to 27 minutes.
Every minute spent waiting translates directly to increased latency for our users or forces us to keep GPUs active "just in case" to avoid the wait.
Boot time also taxes everything else we do: deployment rollouts, RL training jobs that spin up inference servers, and research iteration speed all sit behind the same cold-start wall.
To enable effective autoscaling, we distinguish between two startup scenarios: Cold start: Occurs when we deploy a model configuration for the first time.
Warm start: Occurs when we scale up an existing deployment or during the consecutive phases of a rolling update.
…





