EdgeCampus runs a shared GPU server for student research jobs: 4 GPUs, 32 CPU cores, 256 GB RAM. Jobs arrive in a queue; each job needs 1 GPU, 4 CPU cores and 24 GB RAM, and runs 10–60 minutes. Some students submit tiny test jobs (2 minutes) that wait behind long ones.
The administrator asks: "How many jobs should run at the same time, and how should the waiting jobs be organized?"
- Identify the resources and classify each (preemptable? sharable? reusable/consumable?).
- Compute the maximum number of concurrent jobs allowed by each resource and the resulting bottleneck.
- Explain what goes wrong if the system admits more jobs than the bottleneck allows (think about memory, context switching, GPU contention).
- Propose a design: processes or containers per job, a limit (admission control), and a queue policy that treats the tiny test jobs fairly. Justify with concepts from this chapter and preview which chapters (4, 6, 10, 12) will refine it.
- Define two metrics you would measure for one week to evaluate your design, and the result that would make you change it.