Pingbang Hu
Pingbang Hu 胡平邦
I speak TeX\TeX

I'm a fourth-year Ph.D. candidate at ImageUniversity of Illinois Urbana-Champaign (UIUC) advised by Jiaqi Ma, and I also work closely with Han Zhao.

Previously, I was an ImageAnthropic AI Safety Research Fellow and a Ph.D. ML intern at ImageSusquehanna International Group, and I've also spent time at ImageImageAmazon AWS AI Lab and ImageImageNational Institute of Informatics. I obtained my Master degree from ImageUIUC and dual Bachelor degree from ImageUniversity of Michigan and ImageShanghai Jiao Tong University.

🔬 Research

Research-wise, I'm interested in the broad area of ML and AI, with the goal being to draw theoretical insights from practical problems and develop algorithms with provable guarantees and desirable properties such as efficiency, robustness, and fairness. Recently, my research focuses on understanding data, including the following three aspects:

  1. Data Attribution: Understanding how training data influences AI models.
  2. Data Curation: How to curate/generate/augment (synthetic) data that further helps models generalize?
  3. Data-Centric Privacy: Can above be done without compromising privacy when safety-critical or sensitive data is involved? This includes (differential) privacy, machine unlearning, etc.

Previously I have worked on graph neural networks with Jiaqi Ma and fast graph algorithms with Thatchaphol Saranurak. Generally speaking, I held (actually hold) a strong interest in theoretical stuffs that involves geometry.

🗞️ News

  • Aug 2026

    🎤 Giving a talk on Dr. Post-Training at ImageImageGoogle DeepMind!

  • Aug 2026

    🎤 Giving a talk on Towards Market Data Valuation under Complex Training at ImageSIG!

  • Jun 2026

    💼 Interning at ImageSIG Deep Learning team, come hanging out in Philly!

  • May 2026

    🎤 Giving a talk on Science of Data: Predictable, Optimizable, and Scalable at ImageImageCitadel GQS!

  • May 2026

    🎤 Giving a talk on Agentic Backdoor via Pre-Training Poisoning at ImageAnthropic!

  • Feb 2026

    🏛️ We are organizing the Data Foundations of AI, come check out if you work on data as well!

  • Jan 2026

    📝 One paper accepted by ICLR 2026.

    PNL
  • Jan 2026

    💼 Starting as an AI safety fellow at ImageAnthropic, come hanging out in San Francisco!

  • Oct 2025

    🏛️ We are organizing the Symposium on Information Retrieval and Language Models at ImageUIUC!

  • Oct 2025

    🎤 Giving a tutorial on recent tricks in computing gradient-based data attribution!

    GraSS
  • Sep 2025

    📝 Please check out our new survey paper on data attribution!

    Survey
  • Sep 2025

    📝 Two papers accepted by NeurIPS 2025!

    GraSS · Unlearning
  • Aug 2025

    🎓 Get my M.S. Applied Math Degree at ImageUIUC!

  • Jul 2025

    🏛️ We are organizing the 3rd Workshop on Regulatable Machine Learning in conjunction with NeurIPS 2025!

  • Jul 2025

    🎤 Giving a talk on Data Attribution at the Guided Generation Group (GGG)!

    Slide
  • Jun 2025

    ✈️ Attending the first AI Startup School held by Y Combinator, see you in San Francisco!

  • Mar 2025

    💼 Interning at ImageImageAmazon AWS AI Deep Engine Science team, come hanging out in New York!

  • Jan 2025

    📝 One paper accepted by ICLR 2025.

    Adversarial DA
  • Nov 2024

    🏆 Received the Graduate Conference Travel Award from ImageUIUC!

  • Oct 2024

    🏆 Received the NeurIPS 2024 Scholar Award, see you in Vancouver!

  • Sep 2024

    📝 Two papers accepted by NeurIPS 2024 with one Spotlight.

    dattri · MISS
  • Jun 2024

    📖 We launched the ongoing Data Attribution Reading Group.

  • May 2024

    💼 Interning at ImageImageNII, come hanging out in Tokyo!

📄 Selected Research

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
May 8, 2026
arXiv preprint

We reconceptualize general data in LLM post-training as a data-induced regularizer, unifying existing data selection methods along a bias-variance spectrum.

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
A Unified Theory of Random Projection for Influence Functions
Feb 11, 2026
ICLR 2026 Workshop DATA-FM

A unified theory of random projection for influence functions.

A Unified Theory of Random Projection for Influence Functions
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
Sep 18, 2025
NeurIPS 2025

We propose an efficient gradient compression algorithm to accelerate and scale gradient-based data attribution methods to billion-scale models.

GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
Most Influential Subset Selection: Challenges, Promises, and Beyond
Sep 25, 2024
NeurIPS 2024

We provide a comprehensive study of the common practices in the Most Influential Subset Selection (MISS) problem.

Most Influential Subset Selection: Challenges, Promises, and Beyond

🔖 Misc

I'm from Taiwan 🇹🇼! In my spare time, I enjoy street photography 📷 and playing drums 🥁.