THINK FIRST·CODE LATER

← All labs

Research mini-project: memory-aware placement of AI inference

Problem

EdgeCampus runs three kinds of inference tasks: object detection (model 0.6 GB, 15 ms on edge GPU), speech recognition (model 1.5 GB, 40 ms), and a language assistant (model 6 GB at 16-bit or 1.5 GB at 4-bit, 120 ms on edge, 60 ms in the cloud plus 45 ms network round-trip). The edge server has 12 GB usable GPU memory; loading a model from SSD takes 1 s per GB. Requests arrive at 50/s (detection), 10/s (speech) and 2/s (assistant).

  1. Formulate the placement problem: which models are kept resident on the edge, which requests are offloaded to the cloud, and at what precision? State the objective (e.g., average or 95th-percentile latency) and the constraints (memory, GPU time, accuracy).
  2. Evaluate at least three configurations numerically (latency per task type and weighted average), including one that swaps models in and out on demand. Show your calculations.
  3. Which configuration do you recommend? What changes if the assistant load increases to 20/s, or the network round-trip rises to 150 ms?
  4. Connect your analysis to the literature: this is a joint caching (which models to keep) + offloading (where to run) problem. Name two ideas from research on edge caching or service placement that could improve your solution.

Work it out on paper, in a document or here, then compare with the model answer. Your answer stays in your browser — it is never sent to or stored on the server.