Simulate distributed training jobs on a GPU cluster.
Input: policy (partial or gang), number of GPUs, number of jobs, then per job name workers duration, then the list of pod requests (job names, one token per pod) — all arriving at time 0 in that order.
Events happen at integer times. At each event time:
- partial: go through the waiting pods in order and give each one a free GPU if there is one (pods that get no GPU keep waiting, in order).
- gang: go through the jobs in input order and give a job all its missing GPUs if enough are free (first fit — later, smaller jobs may start before earlier, bigger ones).
- A job whose pods all have GPUs starts now (print
t=0: JA starts on 6 GPUs, jobs in input order) and runs for its duration. - If no job is running at this point, print
t=T: DEADLOCK, holding: JA=4/6 JB=4/6, free GPUs 0(unfinished jobs, input order) and stop. - Otherwise advance to the next finish time; print
t=10: JA finishes, frees 6 GPUs(input order) and release them; repeat.
When all jobs are done print All jobs finished, makespan T.