# Krea — Bending the Kubernetes scheduler

- Company: Krea (krea.ai)
- Announced: 2026-08-07T15:00:00+00:00
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.krea.ai/blog/bending-the-kubernetes-scheduler
- Record: https://forck.live/items/15099-bending-the-kubernetes-scheduler
- Subject: Krea 2 / AI creative platform

The previous post covered scaling inference anywhere. This one covers the other half: how we bend the default Kubernetes scheduler to keep GPU workloads in the cluster and spill only when it's full, using dynamic taints driven by a PromQL signal, the Descheduler, and Kueue. Introduction In the previous post we showed how we built a system that lets us scale anywhere, using the Virtual Kubelet project. But we skipped a big chunk of that system: how we schedule around the VK node to maximize cluster usage and minimize cost. Some approaches we have seen modify the Kubernetes scheduler, by changing its configuration or by adding plugins as steps. Others replace it entirely with schedulers like Volcano , Apache YuniKorn , or KAI-Scheduler . None of these capture our business needs out of the box, so we took a different route: make the default Kubernetes scheduler work for us. We could have written plugins or modified the default scheduler ourselves. During the project, though, we noticed we did not need to. With a bit of cleverness and a few tools built on top, the default scheduler could still do the job. …

---

Record: https://forck.live/items/15099-bending-the-kubernetes-scheduler
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
