Re-Identification in Differentially Private Incomplete Datasets
TL;DR: In this paper , the authors proposed an algorithm for estimating the number of people in a population who have certain attribute values based on incomplete, differentially private databases and concluded that the re-identification risk must be evaluated even after applying state-of-the-art techniques to protect privacy.
read more
Abstract: Efforts to counter COVID-19 reaffirmed the importance of rich medical, behavioral, and sociological data. To make data available to many researchers who can conduct statistical analyses and machine learning, personally identifiable information must be excluded to protect individual privacy. It is essential to remove explicit identifiers, sample population data, and apply differential privacy, the de facto standard privacy metric. Despite the general belief that the risk of re-identification is insignificant when these techniques are applied, this study shows that even after applying these techniques, the risk of being re-identified is highly significant for some data. This study proposes in detail an algorithm for estimating the number of people in a population who have certain attribute values based on incomplete, differentially private databases. If the estimated number is one, the probability that only one person with that attribute value is present in the population is high, which means that there is a high probability of re-identification. Therefore, this study concludes that the re-identification risk must be evaluated even after applying state-of-the-art techniques to protect privacy.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Machine Learning Model Generation With Copula-Based Synthetic Dataset for Local Differentially Private Numerical Data
01 Jan 2022
TL;DR: In this article , the authors focus on the decision tree machine learning algorithm, and instead of applying it as is, they use a preprocessing technique wherein pseudodata are generated using a copula while removing the effect of noise added by differential privacy.
12
A Computational Framework for Preserving Privacy and Maintaining Utility of Geographically Aggregated Data: A Stochastic Spatial Optimization Approach
Yue Lin,Ningchuan Xiao +1 more
TL;DR: In this article , the authors developed a methodological framework to address the issues of privacy protection for geographically aggregated data while preserving data utility, where individuals at high risk of disclosure are moved to other locations to protect their privacy.
6
Data collection of biomedical data and sensing information in smart rooms
Yuichi Sei,Akihiko Ohsuga +1 more
TL;DR: In this paper , the authors presented a new dataset, including behavioral, biometric, and environmental data, obtained from 23 subjects each spending 1 week to 2 months in smart rooms in Tokyo, Japan.
3
Local Differential Privacy for Person-to-Person Interactions
01 Jan 2022
TL;DR: Li et al. as discussed by the authors proposed a mechanism that satisfies LDP in a person-to-person interaction scenario, which can maintain high data utility while ensuring LDP compared to existing methods.
2
References
k -anonymity: a model for protecting privacy
TL;DR: The solution provided in this paper includes a formal protection model named k-anonymity and a set of accompanying policies for deployment and examines re-identification attacks that can be realized on releases that adhere to k- anonymity unless accompanying policies are respected.
9.2K
On the Lambert W function
TL;DR: A new discussion of the complex branches of W, an asymptotic expansion valid for all branches, an efficient numerical procedure for evaluating the function to arbitrary precision, and a method for the symbolic integration of expressions containing W are presented.
6.4K
L-diversity: privacy beyond k-anonymity
Ashwin Machanavajjhala,Johannes Gehrke,Daniel Kifer,Muthuramakrishnan Venkitasubramaniam +3 more
- 03 Apr 2006
TL;DR: This paper shows with two simple attacks that a \kappa-anonymized dataset has some subtle, but severe privacy problems, and proposes a novel and powerful privacy definition called \ell-diversity, which is practical and can be implemented efficiently.
RAPPOR: Randomized Aggregatable Privacy-Preserving Ordinal Response
TL;DR: RAPPOR as discussed by the authors is a system for crowdsourcing statistics from end-user client software, anonymously, with strong privacy guarantees, allowing the forest of client data to be studied, without permitting the possibility of looking at individual trees.
2K