A scheduling problem that batch and grid computing systems solved in the 1990s has spent most of Kubernetes’ history outside Kubernetes. Gang scheduling — the requirement that a set of coordinated tasks start together or not at all — was never something the default kube-scheduler could express. The workaround was to run a second scheduler alongside the first. With v1.36, that workaround starts to become unnecessary.

One correction to the framing, because the version numbers matter for anyone planning an upgrade: gang scheduling did not first arrive in v1.36. It landed in v1.35, which introduced the foundational Workload API, basic gang scheduling, and opportunistic batching for identical pods. What v1.36 does is restructure that foundation into something the scheduler can build on — and, in the process, move the coordination logic into the default scheduler’s own scheduling cycle rather than bolting it to the side.

The Coscheduling Problem, Restated

The theory here is old and well understood. A distributed training job running across 64 GPUs is a single synchronized computation: every participating process must be running before any useful work happens. Schedule 63 of them and you have not achieved 98% of the job. You have achieved nothing, while holding 63 GPUs idle and blocking every other job that could have used them. John Ousterhout’s coscheduling work in the early 1980s and the gang-scheduling literature that followed established the basic result — that for tightly coupled parallel tasks, all-or-nothing placement is not an optimization but a correctness requirement for making progress.

Kubernetes’ default scheduler was built on the opposite assumption. It places one pod at a time, and it is very good at that. A stateless HTTP service does not care whether its siblings are running. The scheduler’s per-pod model was a deliberate simplification that matched the canonical cloud-native workload of the mid-2010s, and it is the same simplification that made Kubernetes’ first decade as an orchestrator possible.

AI workloads broke the assumption, and the ecosystem responded with secondary schedulers. NVIDIA’s KAI Scheduler, Volcano, and the earlier kube-batch all implement gang semantics by running as a separate scheduler process for batch workloads while kube-scheduler handles everything else. As documented in the research behind Kubernetes GPU scheduling, that pattern works, and in production it works well.

It also has a structural defect. Two schedulers making independent placement decisions over one shared resource pool have no common view of what has been committed. Each can believe a node has capacity that the other has already claimed. The usual mitigations — partitioning nodes between schedulers, or accepting occasional binding conflicts and retries — trade either utilization or predictability for the gang guarantee. That is a reasonable trade to make, but it is a trade forced by the architecture rather than by the problem.

What v1.36 Actually Changes

Kubernetes v1.36, released as “Haru” on April 22, 2026, ships a suite of Workload-Aware Scheduling features in alpha, detailed in the project’s own deep-dive published May 13. The API group moves to scheduling.k8s.io/v1alpha2, entirely replacing the v1alpha1 group that v1.35 introduced — an upgrade note worth flagging for anyone who adopted the alpha early, since this is a replacement rather than an addition.

The central architectural change is a separation of concerns that v1.35 had conflated. In v1.35, a single Workload resource carried both the static definition of pod groups and their runtime scheduling state. In v1.36, those split:

  • Workload becomes a static template only, declaring spec.podGroupTemplates[] with a scheduling policy per template — for example gang: {minCount: 4}, meaning the group is schedulable only when at least four pods can run simultaneously.
  • PodGroup (KEP-5832) becomes a separate runtime object holding actual scheduling state, referencing its template origin and carrying owner references back to the real workload object such as a Job.

The practical payoff is that the scheduler reads PodGroup directly. It no longer has to watch and parse Workload objects to determine runtime state, which is what makes the design plausible at scale. Per-replica sharding of status updates addresses the write amplification that would otherwise make a single status object a bottleneck for large groups.

On top of that separation, v1.36 adds a PodGroup scheduling cycle inside kube-scheduler — the piece that makes the title’s claim true in the way that matters. Atomic group placement becomes a first-class operation of the default scheduler, not an external one. Alongside it come first iterations of topology-aware scheduling and workload-aware preemption, plus ResourceClaim support that connects PodGroups to Dynamic Resource Allocation, and phase one of native Job controller integration (KEP-5547) built on the gang scheduling foundation of KEP-4671.

Workload-aware preemption deserves particular attention because it fixes the mirror image of the original problem. Gang scheduling ensures you do not partially place a group. Partial preemption is the same defect running backwards: evicting three pods of a four-pod gang to make room for something else destroys the victim job without freeing enough capacity to help the newcomer. The v1.36 mechanism evaluates preemption at group granularity, asking whether evicting entire lower-priority groups would admit a whole higher-priority one.

The Caveat That Governs Adoption

Every one of these features is alpha, and the preemption mechanism is explicitly disabled by default. This is not production-ready scheduling infrastructure, and the v1alpha1-to-v1alpha2 replacement is a reminder that the API surface is still moving underneath adopters.

That tempers the practical advice but not the architectural significance. The secondary-scheduler pattern was never a design anyone chose on its merits; it was the only way to express a constraint the core API could not represent. Once the default scheduler can express gang semantics natively, the question for teams running Volcano or KAI shifts from “which secondary scheduler” to “when does the in-tree path become good enough to retire ours” — and the answer will arrive over several releases, not one.

For anyone evaluating that timeline: watch the Job controller integration phases rather than the scheduling features themselves. Gang scheduling that requires hand-authored Workload and PodGroup objects serves a narrow audience of platform teams. Gang scheduling that a plain Kubernetes Job inherits by declaring a completion count is what moves the capability from specialist infrastructure to a default that most clusters get for free. Phase one shipped in v1.36; the useful question for the next two releases is how much of the Job API surface it eventually covers.