Ankit Pensia
Assistant Professor in the Department of Statistics and Data Science at Carnegie Mellon University.
I develop reliable and computationally efficient methods for high-dimensional inference when data are corrupted, heavy-tailed, missing, or subject to resource constraints. A recurring theme in my work is the gap between what is statistically possible and what can be computed efficiently.
Research Areas and Selected Publications
A high-level overview of my research is in this short talk, given at Simons Institute. Broadly, my research falls into three areas. Use to expand the plain-language takeaways for the selected papers below.
1. Robust and Heavy-Tailed Statistics
Outliers and heavy-tailed distributions pose significant challenges to standard inference procedures and are studied in the field of robust statistics. I'm interested in the statistical and computational landscape of robust algorithms.
- SoS Certifiability of Subgaussian Distributions and its Algorithmic Applications
STOC, 2025
We prove that for each subgaussian distribution, there is a small sum-of-squares proof that its moments are bounded. This result is surprising as it goes against the conventional wisdom of the community. Our result dramatically expands the scope of the current algorithmic toolkit to more general distributions. - A Sub-Quadratic Time Algorithm for Robust Sparse Mean Estimation
ICML, 2024 (Spotlight)
Existing algorithms for robust sparse estimation ran in $\Omega(d^2)$ time, where $d$ is the dimension. Thus, while sparsity improves the statistical performance (in terms of sample complexity), these benefits do not translate into practice because of high computational cost. We developed the first subquadratic algorithm. For an introduction to this problem (and this result), please see the slides of this survey talk. - Outlier Robust Mean Estimation with Subgaussian Rates via Stability
NeurIPS, 2020
We show that recent outlier-robust algorithms also achieve subgaussian confidence intervals for heavy-tailed distributions, and vice-versa. In particular, we identify the "stability" condition as the bridge between these two contamination models. In a recent paper at NeurIPS 2022, we extended these results to robust sparse mean estimation, i.e., when the mean is is sparse.
2. Inference under Constraints
The proliferation of big data has led to distributed inference paradigms such as federated learning, which impose constraints on communication bandwidth, memory, or privacy. My research focuses on understanding the impact of these constraints.
- Simple Binary Hypothesis Testing under Local Differential Privacy and Communication Constraints
COLT, 2023
We characterize the minmax optimal sample complexity of hypothesis testing under local differential privacy and communication constraints, develop instance-optimal algorithms, and show separation between the sample complexity for binary and ternary distributions. In this paper, we build on our recent prior work, where we investigated the cost of only the communication constraints. - Streaming Algorithms for High-Dimensional Robust Statistics
ICML, 2022
We develop the first (computationally-efficient and sample-efficient) streaming algorithm with $o(d^2)$ memory usage for a variety of robust estimation tasks; in fact, our algorithm uses only $\tilde{O}(d)$ space. All the prior robust algorithms needed to store the entire dataset in memory, which leads to quadratic memory. In a recent paper at ICML 2023, we extended it to robust PCA.
3. Machine Learning and Statistics
I have a broad interest in the fields of machine learning and statistics.
- The Sample Complexity of Simple Binary Hypothesis Testing
COLT, 2024
We revisit the problem of simple binary hypothesis testing, a fundamental problem in statistics. Despite being studied for a century, the non-asymptotic rates of this problem were not known (in contrast, the asymptotic error rates were very well-understood). We close this gap and derive tight sample complexity results.
Teaching
Current Course
Past Courses
Prospective Students
If you're interested in working with me, please apply to CMU's Statistics & Data Science Ph.D. program. Please note that admissions are decided by a central committee, not individual faculty.
About
Previously, I was a research fellow at the Simons Institute (UC Berkeley) and a Herman Goldstine Postdoctoral Fellow at IBM Research.
I received my Ph.D. from the Computer Sciences department at UW-Madison in 2023, where I was advised by Po-Ling Loh, Varun Jog, and Ilias Diakonikolas. My dissertation received the Graduate Student Research Award from the CS Department. Before Madison, I spent five memorable years at IIT Kanpur.
Feel free to email me at [email protected] - I'd love to chat if we share interests. When I'm not scribbling on a whiteboard, you'll usually find me walking somewhere scenic, playing squash, or reading a good book.