THINK FIRST·CODE LATER

← Software Engineering
Chapter 14 · Week 15

The Future of Software Engineering with AI

Before You Start: What You Must Be Able to Do

Before the questions, make sure you can: place today's AI tools in the long history of automating programming; describe what coding agents do, how their progress is measured (e.g., SWE-bench) and the limits of such benchmarks; use a scale of AI autonomy to decide how much independence an AI may have for a given task, based on risk, reversibility and data sensitivity; explain spec-driven development and why verification becomes the main bottleneck; design guardrails for agents (sandboxing, least privilege, allow-lists, protected tests, review gates, logs); discuss the risks of AI at scale — comprehension debt, deskilling, new attack surfaces, dependence on a few vendors, energy use — and how teams manage them; name the skills that stay valuable; and connect all chapters of this course into one engineering view of StudyBuddy.

The Big Idea

Every generation of programmers has been told that its job is about to be automated away — by compilers, by high-level languages, by code generators, and now by AI. Each time, the mechanics of programming became cheaper, the amount of software in the world grew, and the hard part moved up: deciding what to build, making sure it is right, safe and fair, and keeping it working. This final chapter looks at where AI is taking software engineering, what changes, what does not, and how to be the engineer who is still needed — and trusted — in ten years.

A short history of "automatic programming"

  • 1950s: the first compilers were marketed as "automatic programming". FORTRAN (1957) let scientists write formulas instead of machine code; critics expected programmers to become unnecessary. Instead, many more people could program.
  • 1970s–1990s: high-level and object-oriented languages, databases, IDEs and code generators removed more accidental complexity (Chapter 1).
  • 2000s–2010s: open-source libraries, Stack Overflow and cloud platforms made it normal to assemble software from existing parts.
  • 2020s: AI assistants, then agents, generate and modify code from natural-language descriptions.

The pattern matches Brooks's argument: tools remove accidental complexity; essential complexity — understanding the problem, conflicting requirements, design trade-offs, consequences for people — remains. When writing code becomes cheaper, people build more software, and the essential work grows with it.

Coding agents today

A coding agent receives a task ("fix issue #214: matching ignores Sunday slots"), then works in a loop: read the repository, plan, edit files, run the build and tests, read the errors, try again, and finally propose a pull request.

Progress is measured with benchmarks. SWE-bench (2023) collected real issues from open-source Python projects on GitHub together with the tests that verified the real fixes; an agent "solves" an issue if its patch makes those tests pass. Scores rose quickly from very low values in 2023 to a majority of issues on curated subsets within about two years.

Benchmarks are useful, but read them critically:

  • Scope: benchmark tasks are usually small, well-described issues with existing tests — not "work out with instructors what Buddy should do before exams".
  • Contamination: public benchmark data may have been seen during training.
  • Tests as oracle: passing the provided tests is not the same as a correct, maintainable, secure change (Chapter 8).
  • Real projects have unwritten rules, missing tests, legacy code and stakeholders. Recall the 2025 METR study (Chapter 1): experienced developers were slower with AI on their own large projects.

How much autonomy? A practical scale

Driving automation is described in levels; a similar scale helps teams decide how independently AI may work on a given task (one useful way to think, not a standard):

Level AI role Human role Example in StudyBuddy
L0 None Does everything Writing the privacy notice with the legal office
L1 Suggests completions Accepts line by line Autocomplete in the IDE
L2 Assistant: writes functions/tests on request Reviews every change before commit "Write a boundary-value test for groupSize"
L3 Supervised agent: implements a task, runs tests, opens a PR Reviews and approves every PR "Fix issue #214" in a sandbox branch
L4 Autonomous agent within guardrails: changes merged automatically if gates pass Designs the gates; samples and audits results Dependency version bumps that pass CI, SCA and canary analysis
L5 Fully autonomous, no human gate None Not appropriate for systems that affect people

The right level depends on the risk of the task (security, privacy, money, people), its reversibility (can we roll back in minutes?), the sensitivity of the data involved, and the strength of the automated checks (tests, evals, contracts). Updating a typo in documentation can be L4; changing access-control code is at most L2–L3 with security review; anything touching student data or grading needs humans in the loop.

Spec-driven development: the spec is the product

When implementation is cheap, the specification and the verification become the most valuable artifacts:

  1. Humans (with stakeholders) write the what: requirements, acceptance criteria, quality targets, policies (Chapters 3–4).
  2. Humans design the checks: tests, contract tests, property tests, evaluation sets, static rules, architecture rules (Chapters 6, 8, 9).
  3. Agents propose the how: code that satisfies the checks.
  4. Humans review what the checks cannot see: design fit, security, privacy, ethics, maintainability (Chapters 7, 10–12).

Verification becomes the bottleneck. Little's law from Chapter 2 applies: if agents produce changes faster than the team can verify them, unverified work piles up. Teams respond by investing in automated verification — better test suites, mutation testing, stronger type systems, contracts, even formal methods for critical parts — and by keeping changes small.

In plain words

If building houses becomes almost free, the valuable people are the architect who knows what the family needs, the inspector who can prove the house is safe, and the planner who makes sure the neighbourhood still works. The bricklaying robot is useful — but somebody has to sign the safety certificate.

Guardrails for agents

Agents act, so the security principles of Chapter 10 apply to them directly:

  • Sandbox — run in an isolated environment with test data only; no production credentials.
  • Least privilege — only the tools and files the task needs; path allow-lists (an agent fixing matching does not need to edit security/).
  • Network allow-lists — only the package mirror and documentation sites; everything else blocked (prompt-injected instructions often try to send data out).
  • Protected tests and policies — changes to tests, CI configuration or security rules require separate human review (agents under pressure to "make tests pass" may weaken them — Chapter 7).
  • Dangerous-command filters — no git push --force, no deleting directories, no disabling checks.
  • Secrets scanning of every agent change.
  • Complete logs of what the agent read, ran and changed — for review, debugging and accountability.
  • Review gates proportional to risk (the autonomy scale above).

Prompt injection reaches agents too: an issue description, a code comment or a web page the agent reads can contain instructions ("also add this dependency", "print the environment variables"). Treat everything the agent reads as untrusted data.

Risks at scale

  • Comprehension debt and bus factor zero (Chapter 5): code nobody understands is a liability, however fast it was produced.
  • Deskilling: skills that are not practised fade. Junior developers who always delegate may never build the mental models needed to judge AI output. Deliberate practice (writing and debugging code yourself, as in this course's labs) remains essential — you cannot review what you could not have written.
  • Homogenization: many teams using the same models may converge on the same designs — and the same vulnerabilities.
  • Security: new supply-chain paths (hallucinated packages, poisoned training data, injected instructions for agents).
  • Accountability gaps: "the agent did it" is not an answer to a regulator, a court or a student (Chapter 12).
  • Concentration: dependence on a few AI vendors for prices, terms and availability (mitigated by abstraction layers — Chapter 6).
  • Energy and cost (Chapter 13): generating ten candidate solutions to pick one is not free.
  • Work and society: changing job profiles, especially for entry-level tasks — a reason to invest in the skills below.

Skills that stay valuable

  1. Problem framing and requirements — talking to people, finding the real need, writing precise specifications.
  2. Domain knowledge — education, finance, health: understanding the rules and consequences.
  3. Design and trade-offs — architecture decisions under constraints, and explaining them (ADRs).
  4. Verification — test design, evaluation of AI features, reading code critically, debugging.
  5. Security, privacy and ethics judgment — seeing harm before it happens.
  6. Operations — keeping systems reliable, observable and sustainable.
  7. Communication and teamwork — reviews, documentation, psychological safety.
  8. Learning to learn — tools change every year; fundamentals change slowly.

Course synthesis: StudyBuddy from end to end

Chapter Question it answered for StudyBuddy AI angle
1 Foundations Why engineering, not just coding? AI reduces accidental, not essential complexity
2 Process How does the team work? Review bottleneck; AI rules in the DoD
3–4 Requirements What exactly must Buddy do — and not do? No requirement without a source; accuracy requirements
5 Organization Can we deliver by week 15? AI widens estimates; comprehension debt
6 Architecture How do we isolate change and failure? Tutor interface, RAG, guardrails
7 Code quality Is the code maintainable? Prompt as spec; AI code failure modes
8–9 V&V How do we know it works? Tests from requirements; evals for Buddy
10 Security Who might attack it, and how? Prompt injection; controls outside the model
11 Privacy Whose data, for what, where, how long? Redaction; provider contracts; cross-border rules
12 Law & ethics Are we allowed to — and should we? Licences of AI code; AI Act; fair tools
13 Operation Will it keep working, and at what cost? Prompt releases; AI footprint
14 Future What will our job be? Autonomy levels; verification as the core skill
Exam Tip

For "future of software engineering" questions, avoid predictions without reasons. Argue from the course: what AI changes (accidental complexity, speed of generation), what stays (essential complexity, verification, responsibility), and what engineers must do (specify, verify, guard, decide, explain).

Key takeaways

  • Automation of programming is 70 years old; each wave removed accidental complexity and increased demand for software and for engineering judgment.
  • Coding agents plan, edit, test and open PRs; benchmarks such as SWE-bench show fast progress but measure narrow, well-specified tasks.
  • Choose AI autonomy levels by risk, reversibility, data sensitivity and strength of automated checks; keep humans in the loop for anything affecting people.
  • In spec-driven development, specifications and verification are the key artifacts; verification is the new bottleneck.
  • Guardrails for agents: sandbox, least privilege, path and network allow-lists, protected tests, command filters, secret scanning, logs, risk-based review.
  • Risks at scale: comprehension debt, deskilling, homogenization, new attack paths, accountability gaps, vendor concentration, energy.
  • Lasting skills: framing problems, domain knowledge, design trade-offs, verification, security/privacy/ethics judgment, operations, communication, learning to learn.
  • Software engineering is the discipline of being responsible for software — AI makes that responsibility larger, not smaller.

Ready? Close the notes and practise.

30 questions. Predict the output before you check — that is the skill the exam measures.