THINK FIRST·CODE LATER

← All labs

GPU cluster: pod-by-pod vs. gang scheduling

Problem

Simulate distributed training jobs on a GPU cluster.

Input: policy (partial or gang), number of GPUs, number of jobs, then per job name workers duration, then the list of pod requests (job names, one token per pod) — all arriving at time 0 in that order.

Events happen at integer times. At each event time:

  • partial: go through the waiting pods in order and give each one a free GPU if there is one (pods that get no GPU keep waiting, in order).
  • gang: go through the jobs in input order and give a job all its missing GPUs if enough are free (first fit — later, smaller jobs may start before earlier, bigger ones).
  • A job whose pods all have GPUs starts now (print t=0: JA starts on 6 GPUs, jobs in input order) and runs for its duration.
  • If no job is running at this point, print t=T: DEADLOCK, holding: JA=4/6 JB=4/6, free GPUs 0 (unfinished jobs, input order) and stop.
  • Otherwise advance to the next finish time; print t=10: JA finishes, frees 6 GPUs (input order) and release them; repeat.

When all jobs are done print All jobs finished, makespan T.

Write it here or in your IDE, then paste it. Compile and test it yourself before comparing. Your code stays in your browser — it is never sent to or stored on the server.