Pingbang Hu
Pingbang Hu 胡平邦
I speak TeX\TeX

I'm a third-year Ph.D. candidate at ImageUniversity of Illinois Urbana-Champaign (UIUC) advised by Jiaqi Ma, and I also work closely with Han Zhao. Currently, I'm an ImageAnthropic AI Safety Research Fellow and a Ph.D. ML intern at ImageSusquehanna International Group.

Previously, I've spent time at ImageImageAmazon AWS AI Lab and ImageImageNational Institute of Informatics. I obtained my Master degree from ImageUIUC and dual Bachelor degree from ImageUniversity of Michigan and ImageShanghai Jiao Tong University.

🔬 Research

Research-wise, I'm interested in the broad area of ML and AI, with the goal being to draw theoretical insights from practical problems and develop algorithms with provable guarantees and desirable properties such as efficiency, robustness, and fairness. Recently, my research focuses on understanding data, including the following three aspects:

  1. Data Attribution: Understanding how training data influences AI models.
  2. Data Curation: How to curate/generate/augment (synthetic) data that further helps models generalize?
  3. Data-Centric Privacy: Can above be done without compromising privacy when safety-critical or sensitive data is involved? This includes (differential) privacy, machine unlearning, etc.

Previously I have worked on graph neural networks with Jiaqi Ma and fast graph algorithms with Thatchaphol Saranurak. Generally speaking, I held (actually hold) a strong interest in theoretical stuffs that involves geometry.

🗞️ News

  • Jun. 2026

    💼 Interning at ImageSIG Deep Learning team, come hanging out in Philly!

  • Feb. 2026

    🏛️ We are organizing the Data Foundations of AI, come check out if you work on data as well!

  • Jan. 2026

    📝 One paper accepted by ICLR 2026.

    PNL
  • Jan. 2026

    💼 Starting as an AI safety fellow at ImageAnthropic, come hanging out in San Francisco!

  • Oct. 2025

    🏛️ We are organizing the Symposium on Information Retrieval and Language Models at ImageUIUC!

  • Oct. 2025

    🎤 Giving a tutorial on recent tricks in computing gradient-based data attribution!

    GraSS
  • Sep. 2025

    📝 Please check out our new survey paper on data attribution!

    Survey
  • Sep. 2025

    📝 Two papers accepted by NeurIPS 2025!

    GraSS · Unlearning
  • Aug. 2025

    🎓 Get my M.S. Applied Math Degree at ImageUIUC!

  • Jul. 2025

    🏛️ We are organizing the 3rd Workshop on Regulatable Machine Learning in conjunction with NeurIPS 2025!

  • Jul. 2025

    🎤 Giving a talk on Data Attribution at the Guided Generation Group (GGG)!

    Slide
  • Jun. 2025

    ✈️ Attending the first AI Startup School held by Y Combinator, see you in San Francisco!

  • Mar. 2025

    💼 Interning at ImageImageAmazon AWS AI Deep Engine Science team, come hanging out in New York!

  • Jan. 2025

    📝 One paper accepted by ICLR 2025.

    Adversarial DA
  • Nov. 2024

    🏆 Received the Graduate Conference Travel Award from ImageUIUC!

  • Oct. 2024

    🏆 Received the NeurIPS 2024 Scholar Award, see you in Vancouver!

  • Sep. 2024

    📝 Two papers accepted by NeurIPS 2024 with one Spotlight.

    dattri · MISS
  • Jun. 2024

    📖 We launched the ongoing Data Attribution Reading Group.

  • May 2024

    💼 Interning at ImageImageNII, come hanging out in Tokyo!

📄 Selected Research

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
May 8, 2026
arXiv preprint

We reconceptualize general data in LLM post-training as a data-induced regularizer, unifying existing data selection methods along a bias-variance spectrum.

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
A Unified Theory of Random Projection for Influence Functions
Feb 11, 2026
ICLR 2026 Workshop DATA-FM

A unified theory of random projection for influence functions.

A Unified Theory of Random Projection for Influence Functions
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
Sep 18, 2025
NeurIPS 2025

We propose an efficient gradient compression algorithm to accelerate and scale gradient-based data attribution methods to billion-scale models.

GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
Most Influential Subset Selection: Challenges, Promises, and Beyond
Sep 25, 2024
NeurIPS 2024

We provide a comprehensive study of the common practices in the Most Influential Subset Selection (MISS) problem.

Most Influential Subset Selection: Challenges, Promises, and Beyond

🔖 Misc

I'm from Taiwan 🇹🇼! In my spare time, I enjoy street photography 📷 and playing drums 🥁.