THINK FIRST·CODE LATER

← Operating Systems
Chapter 12 · Week 13

Virtualization, Containers and Cloud Resource Management

Before You Start: What You Must Be Able to Do

Before the questions, make sure you can: explain how a hypervisor runs unmodified operating systems (trap-and-emulate, hardware support) and what paravirtualization changes; compute the cost of a TLB miss with nested paging; explain overcommitment of CPU and memory; explain how namespaces and cgroups build containers and compare containers with VMs; describe how the Kubernetes scheduler places pods and what requests, limits and QoS classes mean; predict CFS throttling from cpu.max; place VMs with first fit, best fit and their decreasing variants and compute hosts used and power; apply the HPA scaling rule; and analyse serverless cold starts and keep-alive costs.

The Big Idea

The cloud is an operating system for a data center. It does what a kernel does — multiplex hardware among many users, isolate them from each other, allocate resources and schedule work — but its "processes" are virtual machines, containers and functions, and its "CPU" is thousands of servers. The same questions you studied for one machine reappear at scale: How do we share safely (virtualization, containers)? Where do we place work (bin packing, scheduling)? How much do we give each tenant (limits, autoscaling)? And how do we keep costs and energy down?

Virtual machines and hypervisors

A virtual machine (VM) is a complete emulated computer on which a guest OS runs, managed by a hypervisor (virtual machine monitor):

  • Type 1 (bare-metal): runs directly on hardware — VMware ESXi, Xen, Microsoft Hyper-V, KVM (Linux itself becomes the hypervisor). Used by all clouds.
  • Type 2 (hosted): runs as an application on a host OS — VirtualBox, VMware Workstation. Used on desktops.

Trap-and-emulate. The guest kernel runs in user mode (or a less privileged mode). When it executes a privileged instruction (e.g., disabling interrupts, loading a page-table base), the CPU traps to the hypervisor, which emulates the effect on the VM's virtual state. Early x86 had "sensitive but unprivileged" instructions that did not trap, so VMware used binary translation. Since 2005–2006, Intel VT-x and AMD-V add a special guest mode in hardware, making virtualization efficient and simple.

Paravirtualization (Xen): the guest OS is modified to call the hypervisor directly (hypercalls) instead of executing privileged instructions — faster, but requires changing the guest. Today, drivers are paravirtualized (virtio for disks and networks) even in fully virtualized VMs.

Memory virtualization. A guest translates guest-virtual → guest-physical addresses, and the hypervisor translates guest-physical → host-physical. Hardware nested paging (Intel EPT, AMD NPT) walks both tables. With 4-level tables on both sides, a TLB miss can cost up to (4 + 1) × (4 + 1) − 1 = 24 memory references instead of 4 — which is why huge pages matter even more in VMs.

I/O virtualization: emulated devices (slow, compatible), paravirtual devices (virtio), or device pass-through / SR-IOV (the VM uses a slice of a real network card or GPU directly — near-native speed, used for GPUs in the cloud).

Overcommitment. A host with 64 physical cores may run VMs with 256 virtual CPUs in total (4:1 overcommit) because VMs rarely use all their CPUs at once. Memory can be overcommitted with ballooning (a driver in the guest "inflates" to give memory back), page sharing (deduplicating identical pages) and swapping — carefully, because memory pressure hurts much more than CPU contention.

Containers

A container is not a VM: it is a group of ordinary processes on the host kernel, isolated with two kernel features:

  • Namespaces give a private view: PID (own process numbering, PID 1 inside), network (own interfaces and ports), mount (own file-system tree), UTS (hostname), IPC, user (root inside ≠ root outside).
  • cgroups limit and account resources: cpu.max (quota/period), memory.max, I/O and PIDs.

An image is a stack of read-only layers (union file system, copy-on-write — Chapter 10's COW again) plus a thin writable layer per container; ten containers from the same image share its layers on disk and in the page cache.

VM Container
Isolates Whole OS (own kernel) Processes (shared kernel)
Start time Seconds to minutes Milliseconds to seconds
Overhead Guest OS memory (hundreds of MB) Almost none
Security boundary Strong (hypervisor) Weaker (kernel shared)
Density Tens per host Hundreds per host

Clouds often combine them: containers inside VMs, or micro-VMs (AWS Firecracker, used by Lambda) that start in ~125 ms and give VM-level isolation with container-like speed.

Kubernetes: the cluster scheduler

Kubernetes runs containers in pods on a cluster of nodes. For each new pod, the scheduler:

  1. Filters nodes that can host it: enough allocatable CPU and memory for the pod's requests, matching node selectors/affinity, tolerated taints, free ports.
  2. Scores the remaining nodes (e.g., least allocated → spread load, or most allocated → pack tightly, image locality, topology spread) and picks the best.

Each container declares requests (what the scheduler reserves) and limits (the cgroup ceiling). QoS classes follow: Guaranteed (requests = limits), Burstable (requests < limits), BestEffort (none) — under memory pressure, BestEffort pods are evicted first.

CPU limits and throttling. A CPU limit of 0.2 CPU becomes cpu.max = 20000 100000 (20 ms per 100 ms). A single-threaded request needing 50 ms of CPU runs 20 ms, is throttled for 80 ms, runs 20 ms, waits, and finishes at 210 ms instead of 50 ms. With 4 threads and a limit of 1 CPU, 200 ms of CPU work consumes the quota in 25 ms, waits until the next period and finishes at 125 ms instead of 50 ms. Multi-threaded runtimes (JVM, Go, Node's thread pool) are especially hurt; many teams set requests but avoid tight CPU limits for latency-sensitive services.

Placing VMs: bin packing

Placing VMs (or pods) on as few hosts as possible is a bin-packing problem — NP-hard, so clouds use heuristics:

  • First fit (FF): the first host where the VM fits.
  • Best fit (BF): the host where it fits most tightly.
  • First/best fit decreasing (FFD/BFD): sort VMs by size (largest first), then apply FF/BF. FFD uses at most about 11/9 × OPT + 1 bins in one dimension.

Worked example. Hosts with 10 CPUs and 32 GB; VMs (CPU, GB): A (2, 4), B (5, 8), C (4, 4), D (7, 8), E (1, 2), F (3, 4), G (8, 4). Total CPU = 30 → at least 3 hosts.

  • FF (arrival order): H1 = {A, B, E}, H2 = {C, F}, H3 = {D}, H4 = {G} → 4 hosts.
  • FFD (G, D, B, C, F, A, E): H1 = {G, A}, H2 = {D, F}, H3 = {B, C, E} → 3 hosts, all at 100 % CPU.

Energy. A common linear model: P = P_idle + (P_max − P_idle) × utilization, e.g., 100 W idle and 250 W at full load. FF's 4 hosts draw 850 W; FFD's 3 hosts draw 750 W — 100 W saved by switching one host off. Idle servers still draw ~40–60 % of peak power, so consolidation (packing and turning hosts off) saves energy — at the risk of overload if demand grows, and it needs live migration (Chapter 10) to move VMs.

With several resources (CPU, memory, GPU, network), packing is multi-dimensional: a host full on memory but idle on CPU wastes CPU ("stranded resources"). Good heuristics balance dimensions (e.g., sort by the dominant normalized demand, as in the lab).

Autoscaling

The Kubernetes Horizontal Pod Autoscaler computes:

desired replicas = ⌈ current replicas × current utilization / target utilization ⌉

(skipped if the ratio is within 10 % of 1, clamped to [min, max]). Example: target 60 %, 3 pods at 100 % → ⌈3 × 100/60⌉ = 5 pods; then 5 pods at 90 % → ⌈5 × 1.5⌉ = 8. Autoscaling reacts after load changes — pods need seconds to start and nodes minutes — so systems add headroom (target 60 %, not 95 %), scale-down stabilization windows, or predictive scaling from historical patterns.

Serverless and cold starts

With serverless functions (AWS Lambda, Azure Functions, Knative), the platform creates instances on demand and bills per invocation. A request that finds no idle warm instance triggers a cold start: start a micro-VM/container, load the runtime and the code — hundreds of milliseconds to seconds. Platforms keep instances warm for a keep-alive period after each request: longer keep-alive → fewer cold starts but more idle memory. In the lab example, raising the keep-alive from 0.6 s to 10 s reduces cold starts from 5 to 3 of 8 requests (average latency 412.5 → 287.5 ms), but idle warm time grows from 3.2 s to 38.4 s. This is a classic caching trade-off (like page replacement), studied in research on keep-alive and pre-warming policies.

Noisy neighbours. Tenants on the same host compete for caches, memory bandwidth, disk and network even if CPU and memory are limited. Clouds mitigate with dedicated hosts, cache partitioning (Intel CAT) and I/O limits.

Industry spotlight

Google's Borg (the ancestor of Kubernetes) runs hundreds of thousands of jobs per cluster, mixing latency-sensitive services with batch jobs that use leftover resources and are preempted when needed — the same priority and preemption ideas as CPU scheduling (Chapter 4), at data-center scale.

Research spotlight

Energy-aware VM placement and consolidation, multi-resource fairness (e.g., Dominant Resource Fairness), and SLA-aware autoscaling are active research areas. At the edge, resources are scarce and heterogeneous, so placement, scaling and offloading decisions (Chapter 13) must be made jointly and quickly — often with heuristics like FFD or HEFT, or with learning-based policies.

Key takeaways

  • Hypervisors: type 1 (clouds) vs. type 2 (desktops); trap-and-emulate + VT-x/AMD-V; paravirtualization and virtio; nested paging (up to 24 references per TLB miss); SR-IOV pass-through.
  • Containers = processes + namespaces (view) + cgroups (limits); images are COW layers. Lighter than VMs, weaker isolation; micro-VMs combine both.
  • Kubernetes: filter + score; requests for placement, limits for enforcement; QoS classes.
  • CPU limits cause throttling: quota per period; multi-threaded code hits it faster.
  • VM placement = bin packing; FFD/BFD beat FF/BF; linear power model → consolidation saves energy.
  • HPA: desired = ⌈current × utilization / target⌉; keep headroom.
  • Serverless: cold starts vs. keep-alive cost; noisy neighbours in multi-tenant hosts.

Ready? Close the notes and practise.

30 questions. Predict the output before you check — that is the skill the exam measures.