THINK FIRST·CODE LATER

← Software Engineering
Chapter 8 · Week 8

Verification and Validation I: Testing Fundamentals

Before You Start: What You Must Be Able to Do

Before the questions, make sure you can: distinguish error, fault (defect) and failure; describe what a test case and a test oracle are; name the test levels and explain the test pyramid; design black-box tests with equivalence partitioning, boundary value analysis, decision tables and state-transition testing; measure statement and branch coverage and explain why 100% coverage does not prove correctness; write clear JUnit 5 tests with Arrange–Act–Assert, assertEquals, assertThrows and @BeforeEach, following the FIRST properties; choose between stubs, fakes, mocks and spies; practise test-driven development (red–green–refactor); and judge AI-generated tests — recognizing tests that mirror the implementation, have weak oracles or test nothing.

The Big Idea

Edsger Dijkstra observed that testing can show the presence of bugs, but never their absence. You cannot run every possible input, so testing is a sampling problem: choose the few inputs most likely to reveal defects. That choice is a skill, based on the requirement — not on the code, and not on luck. When AI can generate both code and tests in seconds, knowing how to design a good test is what separates real verification from theatre.

Vocabulary: error, fault, failure

Term Meaning StudyBuddy example
Error (mistake) A human action that produces a wrong result Omar misreads the rule "7 days late still counts"
Fault / defect / bug The wrong code (or document) that results if (daysLate >= 7) return 0;
Failure The system behaves differently from what is required, when the fault is executed A student 7 days late gets 0 instead of 30%

A fault may exist for years without causing a failure — until someone submits exactly 7 days late. That is why tests must execute the risky paths.

A test case has: preconditions (state), inputs, and an expected result. The mechanism that decides whether the actual result is right is the test oracle — usually the requirement, sometimes a reference implementation, a known formula, or a property ("the output list is sorted"). No oracle, no test: running code and "seeing that nothing crashed" is not testing.

Test levels and the test pyramid

From the V-model (Chapter 2):

  • Unit tests — one method or class, in isolation, milliseconds each.
  • Integration tests — several units together, or the code with a real database/API.
  • System tests — the whole application, end-to-end, against the system requirements.
  • Acceptance tests — do users/stakeholders accept it? (the Given/When/Then criteria of Chapter 3).

Mike Cohn's test pyramid: many fast unit tests at the bottom, fewer integration tests, few slow end-to-end (UI) tests at the top. An "ice-cream cone" (mostly manual and UI tests) is slow, fragile and expensive.

Black-box test design

Black-box techniques derive tests from the specification, without looking at the code.

1. Equivalence partitioning (EP). Split the input domain into classes that the program should treat the same way; test one representative of each — including invalid classes.

Requirement: "A post title has 5 to 100 characters." Partitions: length < 5 (invalid), 5–100 (valid), > 100 (invalid). Representatives: 2, 50, 150.

2. Boundary value analysis (BVA). Defects cluster at the edges (< vs <=, off-by-one). Test at and next to each boundary: 4, 5, 6 and 99, 100, 101.

In plain words

If a bridge is rated for 10 tonnes, an inspector does not test 3 tonnes, 4 tonnes and 5 tonnes. She tests 9.9, 10 and 10.1. The edges are where the rules change — and where mistakes hide.

3. Decision tables. When the outcome depends on combinations of conditions, list them in a table; each column (rule) becomes a test.

Rule: who may see a hidden post?

Conditions / rules R1 R2 R3 R4
Viewer is the author Y N N N
Viewer is TA/instructor of the course – Y N N
Post is hidden Y Y Y N
Can view? Yes Yes No Yes

4. State-transition testing. When behaviour depends on history, model the states and events. A post: DRAFT → PUBLISHED → HIDDEN → PUBLISHED (restored) or → DELETED. Test every valid transition at least once, and some invalid ones (e.g., restoring a DELETED post must be rejected).

White-box: coverage

White-box techniques look at the code's structure. Coverage measures how much of it the tests execute:

  • Statement coverage — % of statements executed.
  • Branch (decision) coverage — % of branch outcomes executed (each if both true and false).
int badge(int posts, boolean endorsed) {
    int level = 0;
    if (posts >= 10) {
        level = 1;
    }
    if (endorsed) {
        level++;
    }
    return level;
}

One test badge(12, true) executes every statement (100% statement coverage) but only the true side of each if (50% branch coverage). Adding badge(3, false) reaches 100% branch coverage.

Common confusion: 100% coverage means correct

Coverage tells you what was executed, not what was checked. A test with no assertion executes code and verifies nothing. And coverage cannot find missing code: if the requirement says "zero after 7 days" and nobody wrote that if, every line can be covered while the feature is wrong. Coverage is useful for finding untested code, not for proving tested code is right.

Writing good unit tests (JUnit 5)

A test should read like a small specification. The Arrange–Act–Assert pattern:

class LatePenaltyTest {

    @Test
    void sevenDaysLateKeepsThirtyPercent() {
        // Arrange
        int score = 85;
        // Act
        int result = LatePenalty.scoreAfterLatePenalty(score, 7);
        // Assert
        assertEquals(26, result);            // expected first, actual second
    }

    @Test
    void moreThanSevenDaysLateGivesZero() {
        assertEquals(0, LatePenalty.scoreAfterLatePenalty(85, 8));
    }

    @Test
    void negativeScoreIsRejected() {
        assertThrows(IllegalArgumentException.class,
                     () -> LatePenalty.scoreAfterLatePenalty(-1, 0));
    }
}

Good tests are FIRST:

  • Fast — milliseconds, so developers run them constantly.
  • Independent — no test depends on another's side effects or order.
  • Repeatable — same result every time, on every machine (no real network, no current time, no randomness without a fixed seed).
  • Self-validating — pass/fail automatically; no reading of printed output.
  • Timely — written with (or before) the code, not months later.

Name tests after the behaviour (sevenDaysLateKeepsThirtyPercent), not after the method (test1). Use @BeforeEach to build fresh fixtures for every test.

Test doubles

To test a unit alone, replace its collaborators with test doubles (Gerard Meszaros's terms):

Double What it does StudyBuddy example
Stub Returns fixed answers to calls A calendar stub that always says "today is 10 Oct"
Fake A working but simplified implementation FakeTutor returning canned answers; an in-memory repository
Mock Pre-programmed with expectations; verifies that certain calls happened Verify that notifier.send(ta, …) was called once when a post is hidden
Spy A real or fake object that records how it was used Count how many times the AI provider was called (to test the cache)

Doubles make tests fast and repeatable — the Tutor interface from Chapter 6 exists partly for this. But over-mocking is a trap: a test that mocks everything checks only that the code calls what it calls, and breaks at every refactoring.

Test-driven development (TDD)

Kent Beck's red–green–refactor cycle:

  1. Red — write a small test for the next bit of behaviour; run it; see it fail (proving the test can fail).
  2. Green — write the simplest code that makes it pass.
  3. Refactor — clean up code and tests, keeping everything green.

Benefits: every line of code is covered by a test that was seen failing; the design becomes testable by construction; the tests become an executable specification. TDD combines very well with AI assistance: you write the test from the requirement (the oracle), then let the assistant propose code until it passes — the test is the specification the AI must meet.

The AI angle: AI-generated tests

AI assistants are good at writing the boilerplate of tests and at suggesting cases you forgot ("what about an empty list?"). But their tests have characteristic weaknesses:

  1. The oracle problem. If the AI derives expected values by reading the code, the tests assert what the code does, bugs included. Expected values must come from the requirement.
  2. Weak or missing assertions — assertNotNull(result) or no assertion at all; tests that pass whatever the code does.
  3. Happy-path bias — typical values, few boundaries and invalid inputs.
  4. Over-mocking — everything mocked, so the test checks the implementation's call sequence.
  5. Hidden dependencies — tests using the current date, the network or shared static state: flaky, not repeatable.
  6. Swallowing failures — try { … } catch (Exception e) { } around the assertion, or assertTrue(true).

A quick check of any test suite, human or AI-written: break the code on purpose (change > to >=, return a constant) — does at least one test fail? If not, the tests are too weak. Chapter 9 turns this idea into mutation testing.

Exam Tip

For test-design questions, show your method: list the partitions (valid and invalid), the boundaries (on, just below, just above), and give each test an expected result taken from the requirement. A list of inputs without expected results is not a test design.

Key takeaways

  • Testing shows the presence of bugs, not their absence: choose inputs wisely.
  • Error → fault → failure; every test needs an oracle (expected result).
  • Levels: unit, integration, system, acceptance; keep a pyramid (many fast unit tests).
  • Black-box design: equivalence partitioning, boundary values, decision tables, state transitions — including invalid inputs and transitions.
  • Coverage (statement < branch) finds untested code; 100% coverage ≠ correct.
  • Good unit tests: Arrange–Act–Assert, behaviour names, FIRST, assertThrows for errors.
  • Test doubles (stub, fake, mock, spy) isolate units; avoid over-mocking.
  • TDD: red → green → refactor; your test is the specification an AI must meet.
  • AI-generated tests: check oracles, assertions, boundaries and independence; break the code to test the tests.

Ready? Close the notes and practise.

30 questions. Predict the output before you check — that is the skill the exam measures.