Paraphrasing from David Barstow's super motivating @UCBEPIC keynote: As computer scientists, "is the work you're doing supporting the world of truth? Or is it supporting the world of lies?"
I study humans, AI systems, and humans+AI systems as a CS/HCI Prof @ UCLA. PhD ‘25 @ Berkeley. Previously MIT, EtherPad, Google, Stanford.
- 🎉New Preprint! 🎉 @sh_reya led this investigation of challenges in aligning LLM outputs: *people* don't always have clear criteria, and these "squishy" criteria also drift as users tweak prompts + evaluate outputs. Human input critical in guiding co-evolution of prompts & evals!Evals are arguably the hardest part of LLMOps. LLMs mess up, so we check them w/ other LLMs, but this feels icky. Who validates the validators?? We built an interface to align LLM-based evals with user preferences, learning a lot about why this is hard: arxiv.org/abs/2404.12272
- LOVE this idea!!Replying to @2plus2make5Finally, this project began as a “paper hackathon”, with 6 of us writing code together for 2 days. It was a fun experiment! We got to know each other (and Trader Joe’s snacks) better (movie theater popcorn = best snack).
- Timezone bugs, whee! I wonder what fraction of commit messages are some variant of "fixing timezone issues"...



