Before the questions, make sure you can: choose a deployment strategy (blue-green, canary/staged rollout, feature flags) and plan a rollback; instrument a system with logs, metrics and traces and monitor the four golden signals; define SLIs and SLOs, distinguish them from SLAs, and compute an error budget; run an incident (detect, triage, mitigate, communicate, resolve) and write a blameless postmortem; classify maintenance as corrective, adaptive, perfective or preventive and explain Lehman's laws; manage technical debt with a register and repayment plan; modernize a legacy system with characterization tests and the strangler-fig pattern; explain the environmental footprint of software and AI, apply energy efficiency, hardware efficiency and carbon awareness, and compute a Software Carbon Intensity (SCI) score; and include accessibility (WCAG's POUR principles) and technical sustainability (handover, bus factor) in the definition of "done".
The first release is the beginning of a system's life, not the end. From that day, real users depend on it, the world around it changes, and every decision has a running cost — in money, in people's time, and in energy. This chapter is about keeping software useful, reliable and changeable for years, and about sustainability in three senses: the planet (energy and hardware), people (accessibility and inclusion) and the software itself (can the next team still maintain it?).
Deploying safely
Chapter 1's Knight Capital and CrowdStrike stories were deployment failures. Modern strategies make releases small, observable and reversible:
| Strategy | How it works | Benefit |
|---|---|---|
| Blue-green | Two identical environments; deploy to the idle one ("green"), test it, then switch traffic from "blue" | Instant switch back if something is wrong |
| Canary / staged rollout | Send the new version to a small share of users (1% → 10% → 50% → 100%), compare its error rate and latency with the old version at each stage | A bad release hurts few users and is stopped early |
| Feature flags | New code is deployed switched off; turned on for chosen users (e.g., one course section) by configuration | Separates deploying from releasing; instant "kill switch" |
| Rollback | Return to the previous version (code and configuration) | Recovery in minutes |
Remember that configuration is code: Buddy's prompt template, model name and feature flags must go through the same review, staging and rollout as Java code.
Database changes need special care: use backward-compatible migrations (add a column first, deploy code that uses it, remove old columns later — "expand and contract") so that rollback remains possible.
Observability: knowing what the system is doing
Three kinds of telemetry:
- Logs — timestamped event records ("post 88 hidden by TA t1"). Structured (JSON) logs are searchable. No personal data or secrets (Chapters 10–11).
- Metrics — numbers over time: requests per second, error rate, p95 latency, CPU, AI-provider calls, cache hit rate.
- Traces — the path of one request across components (app → API → tutor → AI provider), showing where time is spent.
Google's Site Reliability Engineering book recommends watching four golden signals for every user-facing service:
- Latency — how long requests take (percentiles!), separately for successes and errors;
- Traffic — how much demand (requests/s, active users);
- Errors — the rate of failed requests (including "successful" responses with wrong content);
- Saturation — how "full" the service is (CPU, memory, connection pools, AI rate limits).
Alert on symptoms users feel (error rate, latency above the objective), not on every internal detail — too many alerts cause "alert fatigue" and real problems get ignored.
SLIs, SLOs, SLAs and error budgets
- SLI (service level indicator) — a measured quantity: the proportion of board requests that succeed within 800 ms.
- SLO (service level objective) — the target for an SLI over a window: 99.5% over 30 days.
- SLA (service level agreement) — a contract with consequences (refunds, penalties) if a level is not met; SLAs are usually looser than internal SLOs.
An SLO implies an error budget: the amount of unreliability allowed. With a 99.5% availability SLO over 30 days (43,200 minutes), the budget is 0.5% = 216 minutes of downtime per month. The budget turns reliability into a shared decision: while budget remains, the team can ship features and take risks; when it is spent, the team freezes risky releases and works on reliability until the budget recovers.
An error budget is like a monthly phone-data plan. You can use your 216 minutes however you like — a risky release, planned maintenance — but when the plan is used up, you stop streaming videos until next month.
100% is the wrong target: it is impossibly expensive, and users cannot tell the difference between 99.99% and 100% because their own Wi-Fi fails more often.
When things go wrong: incident response
A clear process beats heroics:
- Detect — alerts on SLO symptoms, or user reports.
- Triage — severity (e.g., SEV1: Buddy and board down for everyone during exam week; SEV3: one icon missing), assign an incident commander who coordinates and a person who communicates.
- Mitigate first — restore service (rollback, switch the feature flag off, fail over) before searching for the root cause.
- Communicate — status page/notice to students and instructors: what is affected, what to do, when the next update comes.
- Resolve and recover — fix, verify, clean up data.
- Learn — a blameless postmortem.
Blameless postmortems assume that people acted reasonably with the information and tools they had; the question is not "who made the mistake?" but "how did our system make this mistake easy and its detection hard?". A postmortem contains: summary and impact, timeline, root causes and contributing factors (often several), what went well, what went badly, and action items with owners and dates. Blame makes people hide information; blamelessness makes the next incident shorter.
Useful measures: time to detect (MTTD) and time to restore (MTTR, one of the DORA metrics from Chapter 2).
Maintenance and evolution
Maintenance is usually classified (ISO/IEC 14764) as:
| Type | Purpose | StudyBuddy example |
|---|---|---|
| Corrective | Fix defects | Matching ignores Sunday slots |
| Adaptive | Adapt to a changed environment | New Android version, new university SSO, new AI-provider API, new law |
| Perfective | Improve or add functionality/performance for users | Faster search, a better Buddy prompt |
| Preventive | Prevent future problems | Refactoring, upgrading libraries before end of life, adding tests |
Studies have long found that corrective work is only a minority of maintenance; most effort goes to adapting and improving systems.
Lehman's laws of software evolution (Meir Lehman, from the 1970s on, for systems used in the real world) include:
- Continuing change — a system that is used must be continually adapted, or it becomes progressively less satisfactory.
- Increasing complexity — as a system evolves, its complexity increases unless work is done to maintain or reduce it.
- Declining quality — quality will appear to decline unless the system is rigorously maintained and adapted to its environment.
In short: software rots unless someone invests in it.
Technical debt
Ward Cunningham introduced the technical debt metaphor (1992): taking a shortcut now is like borrowing money — you move faster today, but you pay interest (extra effort on every later change) until you repay the principal (clean up the shortcut). Martin Fowler's quadrant distinguishes debt that is deliberate or inadvertent, and prudent or reckless:
| Reckless | Prudent | |
|---|---|---|
| Deliberate | "We don't have time for design." | "We must ship for the midterm; we'll refactor matching in Sprint 6." |
| Inadvertent | "What's layering?" | "Now we know how we should have designed it." |
Managing debt:
- Keep a debt register: item, principal (hours to fix), interest (hours lost per Sprint), risk.
- Pay the items with the best return first (high interest, low principal); reserve a share of every Sprint (often 10–20%) for debt and preventive maintenance.
- Make debt visible to the Product Owner in business terms ("each new notification channel costs 3 extra days until we refactor").
AI-generated code nobody understands (Chapter 5's comprehension debt) is a new, fast-growing kind of technical debt.
Legacy systems and modernization
A legacy system is one that is still valuable but hard to change — old technology, missing tests, missing documentation, original developers gone. Michael Feathers defines legacy code simply as code without tests. Approaches:
- Characterization tests first: tests that record what the system currently does (even if strange), so that changes that alter behaviour are detected.
- Strangler fig pattern (Martin Fowler): build new functionality around the old system and route requests to it piece by piece, until the old system can be switched off — instead of a risky "big-bang rewrite".
- Replace, re-host, re-engineer or retire — choose based on business value and technical condition.
AI assistants are useful here: explaining unfamiliar old code, drafting documentation and characterization tests, and translating code (e.g., COBOL to Java). The risk is the same as in Chapter 7: plausible translations with subtle behaviour changes — which is exactly why characterization tests must exist before any AI-assisted migration.
Sustainable computing 1: the environment
Software has a physical footprint: electricity to run servers, networks and devices, and the embodied emissions of manufacturing hardware. The International Energy Agency estimated that data centres used about 1.5% of the world's electricity in 2024, and projected that this could roughly double by 2030, driven largely by AI. Individual choices by software engineers — multiplied by millions of users — matter.
The Green Software Foundation describes three levers:
- Energy efficiency — do the same work with less energy: efficient algorithms and queries (Chapter 6's caching), no wasteful polling, smaller images and payloads, right-sized servers, idle environments switched off.
- Hardware efficiency — use less hardware, for longer: high utilization instead of many idle servers; keep apps working on older phones (software updates that make old devices unusable create electronic waste).
- Carbon awareness — do flexible work when and where electricity is cleaner (e.g., run nightly index rebuilds when the grid's carbon intensity is lowest).
The Software Carbon Intensity (SCI) specification (standardized as ISO/IEC 21031:2024) gives a rate per functional unit:
SCI = ((E × I) + M) per R
- E — energy used by the software (kWh);
- I — carbon intensity of the electricity (gCO₂e per kWh), which depends on region and time;
- M — embodied emissions of the hardware, allocated to this software;
- R — the functional unit, e.g., per answered question, per user per day.
Because SCI is a rate, it rewards making each unit of work cleaner, not just having fewer users.
The AI footprint. Generative AI is energy-intensive: training large models takes enormous amounts of electricity, and every answer (inference) uses energy on specialized hardware; estimates per query vary widely with model size, answer length and data centre. Engineering choices for Buddy: use the smallest model that meets the quality requirement (measured with the golden set, Chapter 9); cache repeated answers; limit answer length; retrieve fewer, better passages; do not call the model when a simple lookup (FAQ, search) answers the question.
Sustainable computing 2: people — accessibility and inclusion
Software that some people cannot use is not sustainable either. The Web Content Accessibility Guidelines (WCAG 2.2) are organized around four principles — POUR:
- Perceivable — text alternatives for images, captions for videos, sufficient colour contrast (at least 4.5:1 for normal text at level AA).
- Operable — everything works with a keyboard; enough time; no flashing content that can cause seizures.
- Understandable — clear language, predictable navigation, helpful error messages.
- Robust — works with assistive technologies such as screen readers (correct semantic markup).
Accessibility is increasingly a legal requirement too (for example, the European Accessibility Act applies to many digital services from June 2025, and US public institutions follow Section 508 and the ADA). Inclusion also means low-bandwidth and low-end devices (students on mobile data), bilingual interfaces, and not assuming everyone has the newest phone.
Sustainable computing 3: the software itself
Technical and economic sustainability means the system can be maintained affordably by the people who will actually maintain it. StudyBuddy will be handed to next year's student team: that requires readable code, tests, ADRs (Chapter 6), a runbook ("how to deploy, how to roll back, what to do when Buddy is down"), up-to-date dependencies, and a bus factor above 1 (Chapter 5). A system only its original authors can run is a liability.
The AI angle: AI in operations
- AIOps: anomaly detection on metrics, clustering of similar alerts, summaries of long logs and incident timelines — helpful for speed.
- Risks: an assistant that suggests a "fix" during an incident can make things worse; pasting production logs into external tools can leak personal data; automated remediation needs the same safeguards as agents (Chapter 7: least privilege, human approval for impactful actions).
- AI features need operations too: monitor Buddy's quality signals (thumbs-down rate, fallback rate, cost per day, provider latency), version prompts and models, and re-run the evaluation after every change (Chapter 9).
For error-budget questions, always convert the SLO into minutes (or requests) of allowed failure over the window, subtract what was consumed, and state the policy consequence (keep shipping vs. freeze risky changes). For sustainability questions, name the lever (efficiency, hardware, carbon awareness) and a measurable action.
Key takeaways
- Deploy small, observable and reversible: blue-green, canary, feature flags, rollback; configuration is code; backward-compatible migrations.
- Observe with logs, metrics, traces; watch the golden signals (latency, traffic, errors, saturation); alert on user-visible symptoms.
- SLI → SLO → error budget; SLAs are contracts; spend the budget consciously, freeze when it is gone.
- Incidents: detect, triage, mitigate first, communicate, resolve; learn with blameless postmortems and action items.
- Maintenance is corrective, adaptive, perfective, preventive; Lehman: continuing change and increasing complexity.
- Technical debt has principal and interest; keep a register; repay deliberately every Sprint.
- Legacy: characterization tests first, strangler fig instead of big-bang rewrites; AI helps, but verify behaviour.
- Green software: energy efficiency, hardware efficiency, carbon awareness; SCI = ((E × I) + M) per R; right-size AI models, cache, avoid unnecessary calls.
- Accessibility (WCAG POUR) and inclusion are part of quality; plan the handover so the next team can sustain the system.
Ready? Close the notes and practise.
30 questions. Predict the output before you check — that is the skill the exam measures.