Gang Scheduling¶
Gang scheduling ensures all shards of a job start simultaneously or not at all. This is critical for tightly coupled parallel jobs that require synchronization, such as distributed training with collective communication or MPI programs.
What is Gang Scheduling?¶
Gang scheduling is an all-or-nothing scheduling policy:
- Without gang scheduling: Shards start as resources become available
- With gang scheduling: All shards wait until resources for every shard are available, then all start together
Enabling Gang Scheduling¶
Add the --gang flag when submitting sharded jobs:
When to Use Gang Scheduling¶
Required For¶
MPI Programs
MPI jobs require all processes to initialize together:
from mpi4py import MPI
comm = MPI.COMM_WORLD
rank = comm.Get_rank()
size = comm.Get_size()
# All ranks must be present for collective operations
data = comm.gather(local_data, root=0) # Requires all ranks
Distributed Training with Barriers
Training frameworks that use synchronization barriers:
import torch.distributed as dist
# Initialize process group (all ranks must be present)
dist.init_process_group(backend='nccl')
# Barrier synchronization
dist.barrier() # All processes wait here
# Collective operations
dist.all_reduce(tensor) # Requires all processes
Jobs with Collective Communication
Any code using all-reduce, broadcast, scatter/gather, or barriers:
Not Needed For¶
Embarrassingly Parallel Workloads
Independent data processing doesn't need synchronization:
Data Parallelism Without Synchronization
If shards don't communicate, gang scheduling is unnecessary:
Hyperparameter Sweeps
Different experiments running in parallel:
How Gang Scheduling Works¶
Without Gang Scheduling¶
Shards start as soon as resources are available:
Time 0s: Shard 0 starts (GPU available)
Time 0s: Shard 1 starts (GPU available)
Time 30s: Shard 2 starts (GPU became available)
Time 60s: Shard 3 starts (GPU became available)
❌ Problem: Shards 0-1 may timeout waiting for Shards 2-3 to join
With Gang Scheduling¶
All shards wait until all resources are available:
Time 0s: Waiting for 4 GPUs...
Time 0s: 2 GPUs available - not enough, continue waiting
Time 30s: 3 GPUs available - not enough, continue waiting
Time 45s: 4 GPUs available → All shards start simultaneously
✅ Benefit: Synchronized start, no timeouts, no wasted resources
Gang Scheduling and Resource Management¶
Resource Reservation¶
With gang scheduling, resources are reserved atomically:
If only 6 GPUs are available, the job waits rather than starting with partial resources.
Queuing Behavior¶
Gang-scheduled jobs may wait longer but avoid wasted resources:
# This may queue longer but ensures all shards start together
vbatch -N 32 --gang -P cpu-large --name large-simulation python simulate.py
Trade-offs: - Longer queue time: Must wait for all resources - No wasted resources: Avoids partial starts that fail - Guaranteed synchronization: All shards begin together
Combining Gang Scheduling with Other Features¶
With Resource Pools¶
# All shards on A100 GPUs, starting together
vbatch -N 4 --gang -P gpu-a100 --name train python train.py