THINK FIRST·CODE LATER

← All labs

Secure Buddy against prompt injection

Problem

Buddy uses RAG over (a) instructor-uploaded course files and (b) — a new feature request — endorsed answers from the Q&A board, written by students and endorsed by TAs. Its answers are shown in the web and Android apps. The team also wants Buddy to "create a study-group invitation" when a student asks.

  1. Describe four concrete attack scenarios against this design: one direct prompt injection, one indirect injection through course files, one through board answers, and one exploiting the "create invitation" tool. For each, say what the attacker gains.
  2. For each scenario, propose controls outside the model (at least two per scenario) and say where in the request flow they sit (before retrieval, in the prompt, after generation, in the tool layer).
  3. Decide whether the "endorsed board answers" feature should be accepted, and under which conditions.
  4. Write three red-team test cases (input + expected safe behaviour) to add to the golden set (Chapter 9).

Work it out on paper, in a document or here, then compare with the model answer. Your answer stays in your browser — it is never sent to or stored on the server.