Before the questions, make sure you can: explain how privacy differs from security; identify personal data and sensitive personal data; use contextual integrity to judge whether a data flow is appropriate; state the core data-protection principles (lawfulness, fairness and transparency, purpose limitation, minimization, accuracy, storage limitation, security, accountability); describe the main obligations of China's PIPL, the EU's GDPR and the US FERPA as they affect StudyBuddy (without giving legal advice); apply Privacy by Design; build a data inventory with purposes and retention periods; distinguish pseudonymization from anonymization and explain re-identification and k-anonymity; implement data-subject rights; carry out a privacy impact assessment; and design an AI feature that protects personal data in prompts, logs, provider contracts and cross-border transfers.
Security protects data from attackers. Privacy protects people from inappropriate use of their data — including by the organization that legitimately holds it. A perfectly secure StudyBuddy could still violate privacy: by collecting data it does not need, keeping chat histories forever, showing instructors which students "struggle", or sending student questions to an AI provider abroad without telling anyone. Privacy engineering turns legal and ethical principles into concrete design decisions.
This chapter explains principles and laws so that engineers can design responsibly and ask the right questions. It is not legal advice: real projects involve the university's legal and data-protection officers.
Personal data and sensitive data
Personal data (personal information) is any information relating to an identified or identifiable person — directly (name, student ID, email, photo) or indirectly by combining facts (major + dorm + birthday). IP addresses, device identifiers and chat texts that mention the student are usually personal data too.
Some data is sensitive and needs extra protection. China's Personal Information Protection Law (PIPL) lists, for example, biometric data, religious beliefs, specific identities, medical and health information, financial accounts, location tracking, and all personal information of minors under 14. The EU's GDPR has similar "special categories" (health, religion, ethnic origin, sexual orientation, biometrics used for identification…). In StudyBuddy, a student writing to Buddy "I missed the exam because of my depression" has just put health data into a chat log.
Contextual integrity: is this flow appropriate?
Helen Nissenbaum's contextual integrity says privacy is about appropriate flows of information, judged by the norms of the context in which it was shared.
- A student's question to the course Q&A board → classmates and the TA: appropriate (that is the purpose of the board).
- The same question → the student's future employer: inappropriate.
- Buddy chat logs → the instructor to improve the course, aggregated and anonymous: probably appropriate. → the instructor, per named student, "who asked the most basic questions": likely inappropriate and chilling — students would stop asking.
For each data flow ask: who sends, who receives, what data, under which rule (consent, purpose, law) — and would the person reasonably expect it?
Principles of data protection
Most modern laws share principles that go back to the OECD guidelines (1980). The GDPR's version:
| Principle | Meaning | StudyBuddy decision |
|---|---|---|
| Lawfulness, fairness, transparency | A legal basis for processing; no hidden uses; clear notices | A clear privacy notice; consent where required |
| Purpose limitation | Collect for specified purposes; do not reuse for incompatible ones | Availability data is for study matching — not for attendance checks |
| Data minimization | Only what is necessary | No birthday, no home address, no phone number |
| Accuracy | Keep data correct; allow correction | Students can edit their profile |
| Storage limitation | Keep no longer than needed | Chat logs deleted after the semester + appeal period |
| Integrity and confidentiality | Appropriate security | Chapter 10 |
| Accountability | Be able to demonstrate compliance | Data inventory, impact assessment, logs of decisions |
The laws StudyBuddy meets
StudyBuddy is hosted in mainland China, used by WKU students (many Chinese citizens, some international), and connected to Kean University in the United States. Several laws may apply:
- PIPL (China, in force since 1 November 2021) — processing needs a legal basis such as the individual's consent or necessity for a contract or legal duty; separate consent is needed for sensitive personal information, for providing data to other handlers, for public disclosure, and for cross-border transfers. Transfers abroad also need a legal mechanism (a security assessment by the Cyberspace Administration of China for large volumes or important data, a standard contract, or certification) and a personal information protection impact assessment. Individuals have rights to access, copy, correct and delete their data, and a right to an explanation and to refuse decisions made solely by automated decision-making that significantly affect them. Serious violations can be fined up to 50 million yuan or 5% of the previous year's turnover. (China's Cybersecurity Law and Data Security Law add further obligations.)
- GDPR (European Union, since 2018) — applies to organizations outside the EU when they offer services to, or monitor, people in the EU (an exchange student from Germany studying from Europe, for instance). Six legal bases, strong rights (access, rectification, erasure, portability, objection, rights on automated decisions), data-protection impact assessments for high-risk processing, breach notification within 72 hours to the authority, fines up to €20 million or 4% of worldwide annual turnover.
- FERPA (United States, 1974) — protects education records at institutions that receive US federal education funding, such as Kean University: students' records may generally not be disclosed without consent, and students have the right to inspect and correct them. If StudyBuddy data becomes part of Kean's education records, FERPA rules follow.
Laws attach to people and organizations, not only to server locations. Data about EU residents, records of a US university, and transfers of Chinese users' data to a foreign AI provider can each bring other rules into play. Engineers must map whose data flows where — and involve the legal office.
Privacy by Design
Ann Cavoukian's Privacy by Design (2009) has seven foundational principles; the most useful for engineers:
- Proactive, not reactive — prevent problems in design instead of fixing incidents.
- Privacy as the default setting — the user is protected without doing anything: visibility settings start at the most private reasonable option, and optional data (photo, bio, matching) stays off until the user turns it on.
- Embedded into design — part of the architecture, not a plug-in.
- Full functionality (positive-sum) — privacy and useful features, not privacy versus features.
- End-to-end security — for the whole lifecycle of the data, including deletion.
- Visibility and transparency — people can see what happens.
- Respect for user privacy — user-centric choices and controls.
GDPR turned this into a legal duty: data protection by design and by default.
Engineering techniques
1. Data inventory (record of processing). For every data item: purpose, legal basis, source, who can access it, where it is stored (country!), who receives it, retention period.
| Data | Purpose | Access | Location / recipients | Retention |
|---|---|---|---|---|
| Name, WKU email, student ID | Account, notifications | Student, admins | University server (China) | Until graduation + 1 year |
| Enrolled courses | Show boards, matching | Student, staff of course | University server | Semester |
| Posts | Q&A | Course members | University server | Semester + appeal period |
| Buddy chats | Tutoring | Student only | Server; excerpts sent to AI provider | 90 days |
| Free-time slots | Matching | Student; matched classmates see overlaps only | Server | Until changed |
2. Minimization. The cheapest data to protect is data you never collect. Ask for each field: which feature breaks without it? Also minimize what you show (classmates see "free Thursday evening", not the full calendar) and what you send (the AI provider needs the question, not the student ID).
3. Retention and deletion. Define retention per data type and really delete (including backups within their cycle, search indexes and logs). Automated deletion jobs beat good intentions.
4. Pseudonymization vs. anonymization.
- Pseudonymized data replaces direct identifiers with codes (e.g., a keyed hash of the student ID). Whoever holds the key or other data can re-identify people — it is still personal data, but safer to use for analysis.
- Anonymized data can no longer be linked to a person by any reasonably likely means — then data-protection law no longer applies. True anonymization is much harder than it looks.
Re-identification. Latanya Sweeney showed (using 1990 US census data) that about 87% of the US population could be uniquely identified by just ZIP code, gender and date of birth — none of them a "name". In 2008, researchers re-identified users in the "anonymous" Netflix Prize movie-rating dataset by comparing it with public IMDb reviews. In a class of 40, "the only female exchange student from Brazil in CPS 4301" needs no name.
k-anonymity (Sweeney, 2002): a released table is k-anonymous if every combination of quasi-identifiers (age, gender, major, dorm…) appears in at least k records, so each person hides in a group of k. Achieve it by generalizing (age 19 → 18–22) or suppressing rare records. It does not protect against every attack (e.g., if all k people share the same sensitive value), which led to stronger models.
Differential privacy (Cynthia Dwork and colleagues, 2006) adds carefully calibrated random noise to statistics so that the result is almost the same whether or not any single person's data is included. It is used for large-scale statistics (for example by the US Census Bureau for the 2020 census). For StudyBuddy's small classes, a simpler rule is common: do not show aggregate statistics for groups smaller than a threshold (e.g., fewer than 5 students).
5. Rights and transparency. Build features for: seeing and downloading one's data (access, portability), correcting it, deleting it (and its copies), withdrawing consent, and objecting to profiling. Provide a layered privacy notice: a short, clear summary with details one click away, and just-in-time notices where data is used ("Your question will be sent to an external AI service; do not include personal information").
6. Impact assessments. Before high-risk processing — sensitive data, large-scale monitoring, automated decisions, new technologies such as AI, cross-border transfers — carry out a privacy impact assessment (GDPR "DPIA", PIPL "personal information protection impact assessment"): describe the processing, assess necessity and proportionality, identify risks to people, and plan measures.
The AI angle: privacy in AI features
Buddy creates privacy questions that ordinary features do not:
- Personal data in prompts. Students will type names, IDs, health issues and complaints into the chat. Everything sent to the provider leaves the university's control. Controls: a just-in-time warning; automatic detection and redaction of identifiers before sending (emails, phone numbers, ID numbers); never adding the student's identity to the prompt when it is not needed.
- The provider as a recipient. The provider must be contractually bound (a data-processing agreement): no training on StudyBuddy data, limited retention, security measures, location of processing. Check the provider's terms — consumer AI products often have different defaults from business APIs.
- Cross-border transfer. If the AI provider processes data outside mainland China, sending Chinese students' questions to it is a cross-border transfer under PIPL: separate consent, an impact assessment and a legal transfer mechanism may be required. A provider hosted in China, or an on-premises model, may be the simpler design — an architecture decision driven by law (write an ADR!).
- Logs and chat history. Chat logs are among the most sensitive data in the system (students reveal confusion, stress, mistakes). Short retention, access by the student only, no instructor access to individual chats, aggregated analytics only above a minimum group size.
- Inferences and profiling. Analytics that label students as "at risk" are automated profiling with real consequences. Under PIPL and GDPR, people have rights regarding decisions based solely on automated processing; ethically, such labels need transparency, human review and a way to contest them (Chapter 12).
- Memorization and leakage. Models trained or fine-tuned on user data may reproduce it. This is one more reason to use RAG over course materials (Chapter 6) instead of training on student chats.
For privacy questions, structure your answer with the principles: what data (and is it necessary?), for which purpose, who receives it and where, how long, what the person knows and controls (notice, consent, rights), and which assessment is needed. Then name concrete technical measures.
Key takeaways
- Privacy ≠ security: it protects people from inappropriate uses, even by legitimate holders.
- Personal data includes indirect identifiers; sensitive data (health, biometrics, location, minors…) needs extra protection and often separate consent.
- Contextual integrity: judge each flow by the norms of its context.
- Principles: lawfulness/fairness/transparency, purpose limitation, minimization, accuracy, storage limitation, security, accountability.
- PIPL, GDPR and FERPA can all touch StudyBuddy; cross-border transfers and automated decisions need special care; involve the legal office.
- Privacy by Design and by default; build a data inventory with purposes and retention; really delete.
- Pseudonymized ≠ anonymized; quasi-identifiers re-identify people; use k-anonymity, minimum group sizes or differential privacy for released statistics.
- Implement rights (access, correction, deletion, portability) and layered, just-in-time notices; run impact assessments for risky processing.
- AI features: redact personal data from prompts, bind the provider by contract, check cross-border rules, keep chat logs short-lived and private, and be careful with automated profiling.
Ready? Close the notes and practise.
30 questions. Predict the output before you check — that is the skill the exam measures.