Sometimes we had plenty of capacity free on paper, but it was scattered across different machines in chunks too small for a larger job to actually land. Without all-or-nothing admission, the cluster could reach a deadlock, with machines holding resources without doing useful work until someone stepped in manually.
It was clear the status quo wasn’t working and we needed something better.
Choosing the right scheduler
We looked at a few open-source batch schedulers, including Apache YuniKorn, Volcano, and Kueue. YuniKorn didn’t cover all our use cases. And while Volcano had more features, integrating it with our environment would have required replacing core Kubernetes scheduler components.
Kueue won on simplicity and integration. It worked with standard Kubernetes, didn’t require replacing core components, and didn’t force us to rewrite our job specs.
Pairing Kueue with AI Hypercomputer’s flexible, open operations through GKE also contributed to our success. We rely on GKE because it gives us the right level of control for compute-intensive AI work — close access to GPU hardware and drivers, without the overhead of managing raw instances ourselves. Kueue’s native integration with GKE, including with Google Cloud capacity types like Spot VMs and Dynamic Workload Scheduler, meant that once our reserved capacity filled up, the same scheduling logic could reach out to elastic capacity automatically instead of leaving jobs stuck.
Closing the loop with open source
Adopting Kueue turned into something bigger than a simple process change. Because we partner with Google Cloud, we have a direct line to a Technical Account Manager, who saw an opportunity to make AI21 a design partner for the Kueue team. This way, we wouldn’t just be a user, but a source of real production requirements that could help shape where the tool went next.
One of the first things to come out of that partnership was a requirements document we shared with the Kueue team, describing behavior we needed that didn’t exist yet: fair admission ordering across teams for multi-node jobs, without the preemption that usually comes bundled with fairness. The Kueue team built it. Admission Fair Sharing (AFS) reorders the admission queue to favor teams that have historically used less capacity without disrupting jobs already running. Around the same time, we also turned on Topology Aware Scheduling. This existing feature made Kueue aware of our physical cluster layout, so it can refuse to admit jobs that won’t fit on a single node rather than placing them with nowhere to run.
For us, that’s what the best-case open-source feedback loop looks like: Real requirements surfaced through production use, fed directly back into Kueue.
Same fleet, less friction
The Slack #gpu-resources channel is archived now. Every workload, whether a debug pod, a multi-node training run, or an inference deployment, gets queued, prioritized, and scheduled automatically. The results were immediate. Manual interventions dropped from about 20 a week to zero. High-priority jobs that used to wait up to 72 hours now wait 12. Fragmentation across the cluster fell from 15% to 8%, and the “zombie job” problem — workloads partially admitted with nowhere to run — is gone.
None of this changed our total cost. We run our reserved fleet at close to 100% utilization on purpose, so raw spend was never the variable we were optimizing. What changed is where our people’s time goes. Team leads aren’t refereeing compute disputes anymore, and researchers aren’t waiting on replies in Slack. That frees everyone to run more experiments and iterate faster.





