New Ph.D. Student -OR- Visiting PhD Student to work at the intersection of Data Science, Statistics, Privacy & Public Good
Our interdisciplinary team of researchers: Dr. danah boyd (Cornell University), Dr. Rachel Cummings (Columbia University), Dr. Gabriel Kaptchuk (University of Maryland), Dr. Sean Kross (Fred Hutchinson Cancer Center), Dr. Priyanka Nanayakkara (Harvard University), Dr. Elissa Redmiles (Georgetown University), and Dr. Jayshree Sarathy (Northeastern University) is looking for a student or students to conduct application-focused statistical research on data privacy protections, specifically differential privacy. Our team is open to multiple forms of student involvement including:
Differential privacy (DP) is a mathematically rigorous privacy method that involves injecting statistical noise into data. DP has gained significant popularity since its formalization in 2006, and has been recently been deployed by major organizations in industry (eg. Apple, Microsoft, Uber, etc.) and the government (US Census Bureau).
In order for organizations to adopt DP to protect people’s privacy, organizations need a clear understanding of how using DP will impact their work. For example, a medical research organization that wants to protect the privacy of patients whose data they study must ensure their research quality is not diminished by increasing privacy protections. We are working to develop tools that allow data analysts to build a rigorous understanding of how DP may impact the validity of their data analyses. We have two potential projects in mind and look forward to working closely with the candidate to refine these directions to match their interests.
Mathematical analysis methods for propagating error throughout the data analysis pipeline: Widespread deployment of DP is recent, meaning there are few real-world examples of how to mathematically reason about the impact of DP noise on model prediction error rates. This means organizations that want to quantify the relationship between error in training data (e.g., introduced by DP) and robustness of downstream decisions may need to develop their own statistical techniques, as demonstrated by prior work analyzing Facebook’s release of URLs using DP. Moreover, they may be unsure how to integrate other sources of statistical noise (e.g., sampling error) into this analysis in the presence of DP noise. Building on prior work by both Cummings and Sarathy on providing theoretical error bounds when incorporating DP into fundamental AI tasks such as linear regression, mean estimation, and online data analysis, we aim to model end-to-end error and analyze how different sources of noise, including DP, interact and propagate throughout the data analysis pipeline. In addition, we will provide insight into which decisions will benefit from bespoke methods of error analysis versus general methods such as adding up error due to privacy and other sources of noise.
Building software for estimating error introduced by DP: Model builders who are more practically oriented may prefer empirical rather than theoretical metrics of error; additionally, for some AI tasks (such as training neural networks), mathematical analysis of the error may be impractical or impossible. In the absence of privacy considerations, a standard method for empirically estimating a system’s error is simulation, or running the algorithm many times and observing the distribution of errors repeatedly over its output. However, DP methods come with analytical constraints known as composition bounds: when privately training a model on raw data, only a limited amount of model training can be performed before the privacy guarantees are degraded. Intuitively, each DP analysis leaks a small amount of information as controlled by ε, and the values of ε “add up” across multiple analyses, thus limiting the number of AI tasks that can be performed. As a result, model builders cannot freely experiment with DP methods to understand error propagation in the model-to-decision pipeline. Thus, we propose to build a software package that simulates the impact of DP error, which can be distributed with deployed models. The package will create privatized ‘test vectors’ that have the same format as real model releases but different distributions, allowing model builders to simulate the impact of DP error on model outputs and downstream decisions.
The team you will join has a history of work on bringing differential privacy into practice, including analyzing dynamics among interested parties related to adoption, explaining the guarantees of differential privacy to lay people, and developing interactive interfaces for analysts and curators. We have current projects focused on developing and evaluating documentation for DP-protected datasets and eliciting data subjects’ preferences in order to game-theoretically optimize the amount of noise added to datasets by DP in order to balance utility and privacy. Our projects are outward facing, including collaborations with Wikimedia Foundation, Fred Hutchinson Cancer Center, and the Census Bureau. We hope to work closely with the candidate on integrating efforts across projects and supporting their external visibility.
Responsibilities: The researcher will independently design and conduct mathematical and empirical statistical research in collaboration with other members of the research team. The researcher will be responsible for conducting literature reviews, summarizing findings, and publishing results in research journals in collaboration with the research team.
Desired Qualifications:
A strong applicant will have qualifications across all these areas, but the qualifications of each applicant will be considered relative to the specific project interests expressed by the applicant.
Application details vary based on your status. Please review the appropriate section on the next page.
Please email participatoryDP@gmail.com with the subject line “Visiting PhD Student Application” with the following materials by 1/3/25, with applications reviewed on a rolling basis: