The instructor wants a table of study statistics. Before release, check its k-anonymity with respect to the quasi-identifiers age, gender, major.
Input: first line k; then lines age gender major hours (hours = weekly study hours, the sensitive value) until the end of input.
- Group the rows by
(age, gender, major). PrintOriginal: min group size mand, ifm < k,Too small: ...listing each group smaller thankasage/gender/major(size)in order of first appearance, separated by,. - Then generalize the age into 5-year ranges starting at multiples of 5 (
17→15-19,20→20-24) and group again. PrintGeneralized: min group size m, and the small groups in the same way (using the range as age). - Finally, suppress the rows that are still in groups smaller than
k: printSuppressed rows: nandRelease: YES (k=K, rows r)if at least one row remains, otherwiseRelease: NO.
Input:
2
19 F CS 10
20 F CS 12
21 F CS 8
19 M CS 15
22 M CS 5
23 M MATH 9
Output:
Original: min group size 1
Too small: 19/F/CS(1), 20/F/CS(1), 21/F/CS(1), 19/M/CS(1), 22/M/CS(1), 23/M/MATH(1)
Generalized: min group size 1
Too small: 15-19/F/CS(1), 15-19/M/CS(1), 20-24/M/CS(1), 20-24/M/MATH(1)
Suppressed rows: 4
Release: YES (k=2, rows 2)