EdgeCampus runs three kinds of inference tasks: object detection (model 0.6 GB, 15 ms on edge GPU), speech recognition (model 1.5 GB, 40 ms), and a language assistant (model 6 GB at 16-bit or 1.5 GB at 4-bit, 120 ms on edge, 60 ms in the cloud plus 45 ms network round-trip). The edge server has 12 GB usable GPU memory; loading a model from SSD takes 1 s per GB. Requests arrive at 50/s (detection), 10/s (speech) and 2/s (assistant).
- Formulate the placement problem: which models are kept resident on the edge, which requests are offloaded to the cloud, and at what precision? State the objective (e.g., average or 95th-percentile latency) and the constraints (memory, GPU time, accuracy).
- Evaluate at least three configurations numerically (latency per task type and weighted average), including one that swaps models in and out on demand. Show your calculations.
- Which configuration do you recommend? What changes if the assistant load increases to 20/s, or the network round-trip rises to 150 ms?
- Connect your analysis to the literature: this is a joint caching (which models to keep) + offloading (where to run) problem. Name two ideas from research on edge caching or service placement that could improve your solution.