THINK FIRST·CODE LATER

← Software Engineering
Chapter 7 · Week 7

Code Quality and Working with AI Assistants

Before You Start: What You Must Be Able to Do

Before the questions, make sure you can: explain why readability matters more than cleverness; recognize common code smells (long method, duplicated code, magic numbers, long parameter lists, deep nesting, dead code, feature envy, shotgun surgery) and name the refactoring that removes each; refactor in small, behaviour-preserving steps protected by tests; compute McCabe's cyclomatic complexity; explain what static analysis tools find and what they miss; handle errors properly (fail fast, no swallowed exceptions, try-with-resources); write a prompt for an AI assistant that works like a specification (context, constraints, examples, acceptance criteria, request for assumptions and tests); list the typical failure modes of AI-generated code, including hallucinated APIs and packages; review AI-generated code with a checklist; and describe safe practices for coding agents.

The Big Idea

Code is read many more times than it is written — by teammates, reviewers, future maintainers, and you in six months. Quality code is code that is easy to understand, change and test. In 2026 writing code has become cheap: an assistant produces a hundred lines in seconds. What is not cheap is understanding, checking and maintaining those lines. The skills of this chapter — recognizing bad code, improving it safely, and reviewing code you did not write — are now the core of a developer's job.

Readable code

Martin Fowler put it simply: "Any fool can write code that a computer can understand. Good programmers write code that humans can understand."

Names carry most of the meaning:

Poor Better Why
int d; int daysUntilDeadline; Says what it is and its unit
List<Student> list2; List<Student> availableClassmates; Says what it contains
boolean flag; boolean isAnonymous; Booleans read as questions
void process() void hideReportedPost() Methods are verbs that say what they do

Other habits:

  • Small methods doing one thing, at one level of abstraction.
  • Comments explain why, not what. i++; // increment i is noise; // TAs see real names so they can handle misconduct cases (policy MOD-4) is valuable.
  • No magic numbers: if (reports >= 3) → if (reports >= REPORTS_TO_HIDE).
  • Guard clauses instead of deep nesting: handle the special cases first and return.
  • Consistent formatting, done automatically by a formatter — never argue about spaces in a code review.
// Before: nested, magic numbers, unclear names
double f(Post p, int r) {
    if (p != null) {
        if (!p.hidden) {
            if (r >= 3) {
                p.hidden = true;
                return 1;
            }
        }
    }
    return 0;
}

// After: guard clauses, named constant, clear intent
static final int REPORTS_TO_HIDE = 3;

boolean hideIfReportedEnough(Post post, int reportCount) {
    if (post == null || post.isHidden()) {
        return false;
    }
    if (reportCount < REPORTS_TO_HIDE) {
        return false;
    }
    post.hide();
    return true;
}

Code smells

A code smell (Kent Beck and Martin Fowler) is a surface sign that often — not always — indicates a deeper design problem.

Smell Symptom Typical refactoring
Long method 80 lines, many comments separating "sections" Extract method
Large class A class with 40 methods and 25 fields Extract class
Duplicated code The same logic copied in three places Extract method / class; reuse
Magic numbers if (x > 86400) Replace with named constant
Long parameter list match(a, b, c, d, e, f, g) Introduce parameter object
Deep nesting if inside for inside if inside while Guard clauses; extract method
Primitive obsession Times as String "19:00", money as double Small value classes (TimeSlot, Money)
Feature envy A method mostly uses another class's data Move method to that class
Shotgun surgery One small change needs edits in many classes Move related code together
Switch on type The same switch (channel) in many places Replace conditional with polymorphism
Dead code Methods and flags nobody uses (remember Knight Capital!) Delete it — version control remembers
Speculative generality Interfaces and parameters "for the future" that nobody uses Remove until really needed

Refactoring: change the structure, not the behaviour

Refactoring (Fowler) is a change to the internal structure of software to make it easier to understand and cheaper to modify, without changing its observable behaviour.

The discipline:

  1. Make sure tests cover the behaviour you are about to touch (write them first if missing).
  2. Make one small change (rename, extract a method, introduce a constant).
  3. Run the tests. Green → commit. Red → undo and try a smaller step.
  4. Repeat.

Never mix refactoring with new features in the same step: if a test fails, you would not know which change caused it. IDEs perform many refactorings (rename, extract method, inline) automatically and safely — prefer them to manual editing.

Common confusion: "refactoring" = rewriting

Rewriting a module from scratch is not refactoring: it changes behaviour in unknown ways and throws away years of bug fixes. Refactoring moves in small, reversible, tested steps. The same caution applies to "please refactor this file" requests to an AI: check that the behaviour did not change (Chapter 8's tests are your safety net).

Measuring complexity

Cyclomatic complexity (Thomas McCabe, 1976) counts the independent paths through a method: number of decision points + 1. Decision points include if, for, while, case, catch, the ternary ?:, and each && / || in a condition.

int priority(Post p) {                               // 1 (the method itself)
    if (p.isReported() && p.getReports() > 5) {     // +1 (if) +1 (&&)
        return 3;
    }
    for (Tag t : p.getTags()) {                      // +1 (for)
        if (t == Tag.EXAM) {                         // +1 (if)
            return 2;
        }
    }
    return p.isPinned() ? 1 : 0;                     // +1 (?:)
}                                                    // total: 6

A common guideline: keep methods at or below 10; above 20 they are hard to test (you need at least as many test cases as independent paths) and are where bugs cluster.

Static analysis

Static analysis tools examine code without running it: compiler warnings, linters and checkers such as Checkstyle (style), PMD and SpotBugs (bug patterns such as possible null dereferences, unclosed resources, equals without hashCode), and platforms such as SonarQube (smells, complexity, duplication, security hotspots). Run them in CI so every pull request is checked.

What they cannot do: know your requirements. A method that returns a perfectly well-typed wrong answer passes every static check. Static analysis is a net, not a guarantee.

Handling errors properly

  • Fail fast: validate inputs at the boundary and reject bad data immediately with a clear message.
  • Never swallow exceptions:
// Terrible: the error disappears, the post is silently not saved
try {
    repository.save(post);
} catch (Exception e) {
}

// Better: handle what you can, log with context, report what you cannot
try {
    repository.save(post);
} catch (SQLException e) {
    log.error("Saving post {} failed", post.getId(), e);
    throw new PostNotSavedException(post.getId(), e);
}
  • Close resources with try-with-resources (try (Connection c = ...) { ... }).
  • Do not use exceptions for normal control flow, and do not return magic error codes like -1 that callers forget to check.
  • Never log personal data or secrets (passwords, tokens, full chat texts) — Chapter 11.

Working with AI assistants: prompting as specification

An AI assistant produces what you asked for, filling every gap with a guess. A good prompt is therefore a mini-specification (Chapter 4):

  1. Context — language and version (Java 8!), framework, the relevant existing classes and conventions.
  2. Task — one small, clear change.
  3. Constraints — no new libraries, no changes to public interfaces, performance or security rules, style.
  4. Examples / acceptance criteria — input → expected output, including edge cases.
  5. Ask for assumptions — "list every assumption you made".
  6. Ask for tests — but review them yourself (see below).

Weak: "Write a function to match study partners."

Better: "Java 8, no external libraries. Write List<String> matchPartners(Student me, List<Student> others) that returns the names of students enrolled in at least one of my courses whose free slots overlap mine by at least 30 minutes. TimeSlot has LocalTime start, end and DayOfWeek day. Exclude me. Sort by overlap length descending, then name. Examples: … Edge cases: empty lists; slots that touch but do not overlap (19:00–20:00 and 20:00–21:00) do not count. List your assumptions. Do not modify Student."

How AI-generated code fails

AI code usually compiles and looks professional. Its defects are therefore easy to miss. Typical failure modes:

Failure mode Example
Hallucinated APIs Calls list.sortBy(...) or a library method that does not exist
Hallucinated packages Adds a dependency com.fastjson-utils:parser that does not exist — or worse, that an attacker registered after seeing AI tools invent it ("slopsquatting")
Outdated or wrong-version code Java 17 features in a Java 8 project; deprecated APIs
Edge cases Off-by-one errors, empty lists, null, time zones, leap years (Chapter 1)
Insecure defaults SQL built by string concatenation, disabled certificate checks, secrets in code
Silent behaviour changes A "refactoring" that also changes a rounding rule
Plausible but wrong comments // thread-safe on code that is not
Tests that confirm the bug Generated tests assert whatever the code currently returns
Duplication instead of reuse Writes a new date helper instead of using the project's existing one
Over-engineering Factories and interfaces for a 10-line task

Industry analyses of large code bases (for example GitClear's 2024 study of millions of changed lines) reported that, as AI assistance spread, copy-pasted code and code rewritten within two weeks of being written both increased — signs of more churn and less reuse. Speed without review creates maintenance debt.

Remember

The most dangerous AI defect is the one that makes the code look more trustworthy than it is: confident comments, neat formatting, and tests that pass. Review the behaviour against the requirement, not the appearance.

A review checklist for AI-generated code

Before merging AI-assisted code, the committer (and the reviewer) check:

  1. Understand: Can I explain every line? Would I have written something equivalent?
  2. Requirement: Does it do what the acceptance criteria say — no more, no less?
  3. APIs and dependencies: Does every method and package really exist, in the version we use? Is every new dependency necessary, well-known and licence-compatible?
  4. Edge cases: empty, null, zero, negative, very large, boundaries, time zones, concurrency.
  5. Security: user input validated; no string-built SQL; no secrets; access control checked.
  6. Privacy: no personal data logged or sent to new places.
  7. Tests: written or checked by a human from the requirement; they fail when the code is broken (try breaking it!).
  8. Fit: follows project conventions; reuses existing helpers; no speculative complexity.

Coding agents

Agents go further than assistants: they read the repository, run commands, edit many files and propose a pull request. Useful practices:

  • Give them a clear, small task with acceptance criteria — the spec is the prompt.
  • Run them in a sandbox with limited permissions; never give them production credentials or real student data.
  • Review the diff, not the agent's summary of the diff.
  • Require the normal pipeline: tests, static analysis, human review.
  • Watch for agents that "make tests pass" by weakening the tests or special-casing the test inputs.
Exam Tip

If an exam question shows AI-generated code, look systematically: (1) does it compile against the real API/version? (2) edge cases (empty, null, boundaries), (3) security (input, SQL, secrets), (4) does the test actually test the requirement? Name the defect and the fix.

Key takeaways

  • Code is read far more than written: clear names, small methods, guard clauses, comments that explain why, named constants, automatic formatting.
  • Code smells point to design problems; each has standard refactorings.
  • Refactoring changes structure without changing behaviour — in small steps, with tests, never mixed with feature work.
  • Cyclomatic complexity = decisions + 1; keep methods around 10 or less.
  • Static analysis in CI catches bug patterns but cannot check requirements.
  • Handle errors honestly: fail fast, never swallow exceptions, close resources, never log personal data.
  • A good prompt is a specification: context, one small task, constraints, examples, assumptions, tests.
  • AI code fails in plausible ways: hallucinated APIs and packages, wrong versions, edge cases, insecure defaults, tests that confirm bugs, duplication.
  • Review AI code with a checklist; run agents in a sandbox; review the diff; keep humans accountable.

Ready? Close the notes and practise.

30 questions. Predict the output before you check — that is the skill the exam measures.