EdgeCampus wants to run federated/distributed training jobs on three edge sites (8, 8 and 4 GPUs) plus a cloud pool (32 GPUs, 40 ms away). Jobs need between 2 and 16 GPUs, sometimes across sites, and also reserve network bandwidth between the sites they use.
Write a short design proposal (1–2 pages) for a deadlock-free resource manager:
- Identify all the ways jobs could deadlock in this system (GPUs, bandwidth, cross-site reservations, jobs that wait for data from other jobs).
- For each deadlock scenario, choose prevention, avoidance or detection + recovery, and justify the choice (cost, utilization, what information you have in advance).
- Explain how you would combine gang scheduling, resource ordering (e.g., order of sites and resource types), leases and backfilling.
- Define the metrics you would use to evaluate your manager in simulation (e.g., utilization, average job completion time, waiting time of large jobs, number of aborts) and describe one experiment comparing it with pod-by-pod allocation (you may extend Lab OS8.4).
- Discuss one open research question (for example: prediction-based admission using expected rather than maximum demand; fairness between small and large jobs; interaction with offloading latency).