A neural network is a chain of layers. Input: uplinkMbps rttMs inputKB n, then n layers name deviceMs edgeMs outputKB.
Split k (0 ≤ k ≤ n) runs layers 0 … k−1 on the device and k … n−1 on the edge. For k < n, the device sends the data entering layer k (the input for k = 0, else the output of layer k−1) — transfer time = KB × 8 / 1000 / Mbps seconds — then pays the RTT and the edge times. Split n runs everything on the device and sends nothing.
Print split 3 (after pool): send 30 KB, latency 62.0 ms for each k (all on edge for k = 0, all on device for k = n), then Best: split k, latency X ms (the smallest latency; ties → the smaller k).