Before the questions, make sure you can: explain why computation moves to the edge; compute local and remote execution time and device energy for a task; compute an uplink rate with the Shannon formula; derive the break-even bandwidth for offloading; reason about DVFS energy (E ∝ f²); choose a DNN split point between device and edge; explain why selfish offloading decisions can congest an edge server and find a Nash equilibrium by best-response dynamics; use a simple queueing model for edge latency; place tasks with deadlines across device, edge and cloud; and describe how research formulates and solves offloading and scheduling problems.
Every chapter so far asked "which process gets the resource?". Edge computing adds a new question: where should the computation run? A phone can compute locally (slow, drains the battery), send the task to an edge server a few milliseconds away (fast, but shared and limited), or to the cloud (powerful and elastic, but far). The answer depends on the task's size, the data to transmit, the network, the load of the servers, deadlines and energy — and it changes every second as users move. This is computation offloading, a central problem of modern distributed systems and the research area behind this course.
Why the edge?
- Latency: AR/VR needs ~20 ms motion-to-photon; autonomous machines need decisions in milliseconds. A cloud region 30–60 ms away cannot meet such budgets; an edge server 1–5 ms away can.
- Bandwidth: 200 cameras at 2 Mb/s produce 400 Mb/s — process video locally and send results (Chapter 11).
- Privacy and sovereignty: raw data (faces, health data) stays on campus.
- Autonomy: the system keeps working if the internet link fails.
The device–edge–cloud continuum: devices (phones, sensors, cars) → edge (base stations with MEC — multi-access edge computing, standardized by ETSI; campus cloudlets; fog nodes) → cloud data centers. Each layer is more powerful but farther away.
The offloading decision: one task, one user
A task is described by its computation C (CPU cycles) and its input data D (bits); the output is often small.
Local execution on a device running at frequency f_l:
- T_local = C / f_l, E_local = P_compute × T_local.
Remote execution on a server at f_s over an uplink of rate R with round-trip time RTT:
- T_remote = D / R + C / f_s + RTT, E_remote (device) = P_tx × D / R (the phone only pays for transmitting).
Uplink rate (Shannon): R = B × log₂(1 + SNR). With B = 20 MHz and SNR = 10 dB (= 10×): R = 20 × log₂ 11 ≈ 69.2 Mb/s; with B = 5 MHz and SNR = 3 dB: ≈ 7.9 Mb/s.
Worked example — a CampusAR frame: C = 0.2 Gcycles, D = 200 KB (1.6 Mb), phone 2 GHz, P_compute = 0.9 W, P_tx = 1.3 W; edge 20 GHz (GPU-accelerated equivalent), RTT 5 ms; cloud 50 GHz, RTT 60 ms; good Wi-Fi (69.2 Mb/s, upload 23.1 ms).
| Option | Time | Phone energy |
|---|---|---|
| Local | 0.2 / 2 = 100 ms | 0.9 × 0.1 = 90 mJ |
| Edge | 23.1 + 10 + 5 = 38.1 ms | 1.3 × 0.0231 = 30.1 mJ |
| Cloud | 23.1 + 4 + 60 = 87.1 ms | 30.1 mJ |
The edge wins on both time and energy. With weak Wi-Fi (7.9 Mb/s), the upload alone takes 202 ms: local becomes best (100 ms, 90 mJ vs. 217 ms, 263 mJ). A small task with a big input (0.05 Gcycles, 2 MB) should never be offloaded: computing takes 25 ms locally but uploading takes 231 ms.
Break-even bandwidth. Offloading is faster iff D/R + C/f_s + RTT < C/f_l, i.e.
R > D / (C/f_l − C/f_s − RTT)
For the edge: 1.6 Mb / (100 − 10 − 5) ms = 18.8 Mb/s; for the cloud: 1.6 / (100 − 4 − 60) ms = 44.4 Mb/s. If the denominator is ≤ 0 (the task is too small, or the server too far), offloading is never faster. Rule of thumb: offload tasks with high computation per bit of input.
Energy and DVFS. Dynamic CPU energy per cycle grows with the square of the frequency: E_local = κ × C × f². Halving the frequency makes the task 2× slower but uses 4× less energy — so a phone with a loose deadline may prefer to compute slowly rather than offload. Joint decisions on where to run and at which frequency are a classic research formulation.
Partial offloading and DNN partitioning
Many applications are pipelines: split them so the first part runs on the device and the rest on the edge. For a neural network, the best split point balances device computation against the size of the intermediate data (Neurosurgeon, Kang et al., ASPLOS 2017).
Example (input 600 KB; layers conv1, conv2, pool, fc1, fc2 with device times 15, 20, 5, 60, 20 ms, edge times 2, 3, 1, 4, 1 ms and outputs 300, 150, 30, 4, 1 KB; RTT 5 ms):
| Uplink | All on edge | Split after pool (send 30 KB) | All on device | Best |
|---|---|---|---|---|
| 20 Mb/s | 256 ms | 62 ms | 120 ms | split after pool |
| 2 Mb/s | 2,416 ms | 170 ms | 120 ms | all on device |
| 1 Gb/s | 20.8 ms | 50.2 ms | 120 ms | all on edge |
The optimal split moves with the network — a decision the system must re-evaluate at run time.
Many users: congestion and games
An edge server is shared. If each user offloads whenever it is individually better, the server becomes congested and everyone's remote latency increases. Model: the edge's 20 GHz is shared equally among offloading users; each user's remote time = upload 40 ms + RTT 5 ms + C / (20 GHz / k).
With 6 identical users (C = 0.2 Gcycles, local 100 ms): users join one by one while it pays (55, 65, 75, 85, 95 ms…); a 6th user would get 105 ms > 100 ms, so it stays local. Equilibrium: 5 offload (95 ms each), average 95.8 ms. If all offloaded, everyone would get 105 ms — worse than nobody offloading (100 ms)! And the best social outcome is different again: with only 3 offloaders (75 ms each) and 3 local users, the average is 87.5 ms.
This is a game: each user best-responds to the others; the stable outcome is a Nash equilibrium (nobody gains by changing alone). The ratio between the equilibrium's cost and the best possible cost is the price of anarchy. Decentralized offloading games (e.g., Chen et al., IEEE/ACM ToN 2016) prove that best-response dynamics converge, and design mechanisms (pricing, admission control) to improve the outcome.
Queueing view. If an edge server processes μ tasks/s and receives λ tasks/s (M/M/1), the average time in the system is 1 / (μ − λ): with μ = 100/s, λ = 80/s → 50 ms; λ = 95/s → 200 ms. Latency explodes near saturation, so edge schedulers must reject, redirect to the cloud or degrade (lower frame rate/resolution) before the queue grows.
Placement with deadlines and costs
Real tasks have deadlines and resources have prices (cloud time is billed; edge capacity is scarce; phone energy is precious). A simple online policy: process tasks in EDF order and give each one the cheapest resource that still meets its deadline (considering transfer time and when the resource becomes free); if none can, use the resource that finishes it earliest. In the lab example (phone, edge, cloud), 5 of 6 tasks meet their deadlines at a total cost of 2.04 — and the analysis shows an interesting flaw: the hopeless task T5 still occupied the phone and forced T1 to the cloud. Better policies drop or degrade tasks that cannot meet their deadlines.
For workflows (DAGs), offloading becomes the heterogeneous scheduling problem of Chapter 6: device, edge servers and cloud VMs are processors with different speeds and communication costs, and algorithms like HEFT/PEFT extended with energy, cost and deadline constraints are widely used in the literature.
Mobility and orchestration
Users move: their service may migrate between edge servers (Chapter 10's live migration) or keep running remotely at higher latency — a trade-off between migration cost and latency. Edge clusters are managed with lightweight orchestrators (K3s, KubeEdge, OpenYurt) that tolerate unreliable links and small nodes. Federated learning trains models across devices without moving raw data — another placement decision (what to compute where).
This chapter is your instructor's research area. Research on offloading and scheduling in edge–cloud systems typically (1) models tasks (cycles, data, deadlines, DAG dependencies), resources (heterogeneous speeds, energy, prices) and networks (Shannon rates, latency); (2) formulates an objective — minimize latency, energy, cost, or deadline misses — under constraints; (3) proves the problem is hard (often NP-hard, like bin packing or DAG scheduling); (4) designs algorithms: heuristics (list scheduling such as HEFT variants, greedy placement), convex or integer optimization, Lyapunov optimization for long-term averages under uncertainty, game theory for many users, and deep reinforcement learning for dynamic environments; and (5) evaluates them in simulation (e.g., CloudSim, iFogSim, EdgeCloudSim) and on testbeds against baselines. The labs of this chapter are miniature versions of each step.
Telecom operators deploy MEC servers in 5G networks (AWS Wavelength, Azure Operator Nexus); CDNs run code at the edge (Cloudflare Workers, Fastly Compute); cars, drones and factories run local inference with cloud training. The common engineering question is the same as ours: which work must be near the user, and which can go to the cloud?
Key takeaways
- Edge computing: latency, bandwidth, privacy, autonomy; device–edge–cloud continuum (MEC, cloudlets, fog).
- T_local = C/f_l; T_remote = D/R + C/f_s + RTT; device energy local = P_c T_local vs. remote = P_tx D/R; R = B log₂(1 + SNR).
- Offload if R > D / (C/f_l − C/f_s − RTT): high computation per bit, good network, nearby server.
- DVFS: E ∝ C f² — slower can be greener.
- Partial offloading: the best DNN split point depends on bandwidth.
- Many users: congestion, Nash equilibrium, price of anarchy; M/M/1 latency 1/(μ − λ) explodes near saturation.
- Deadline/cost-aware placement (EDF + cheapest feasible resource); DAG offloading = heterogeneous scheduling (HEFT family).
- Mobility → service migration; research uses optimization, Lyapunov, games, heuristics and DRL.
Ready? Close the notes and practise.
30 questions. Predict the output before you check — that is the skill the exam measures.