Log inSign up
Philippe Laban
430 posts
@PhilippeLaban

Philippe Laban

@PhilippeLaban
Research Scientist @MSFTResearch. NLP/HCI Research.
New York City
Joined April 2022
842
Following
1,559
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @PhilippeLaban
    Philippe Laban
    @PhilippeLaban
    Aug 25
    Check out Yoonjoo's fantastic work on simulating users of different expertise levels? It makes progress on this concrete question: for a given task, how do we simulate how a novice vs. an expert would approach it? The validation study is a great example of careful research
    @yoonjoo_le2
    Yoonjoo Lee
    @yoonjoo_le2
    Aug 24
    AI assistants can provide correct, comprehensive information and still be poorly calibrated to the user: too difficult for novices, redundant for experts. KnowSim is a user simulator that tracks what a user knows and updates that knowledge as the conversation unfolds, so we can
    Image
    00:00
  • @PhilippeLaban
    Philippe Laban
    @PhilippeLaban
    Jul 28
    Check out Jihoon's very interesting work on simulating long "situated" conversations where the user can change their mind, switch tasks, underspecify, etc. We need more work evaluating model performance outside of the single-turn, single-task, lab-like setting.
    @jihoontack
    Jihoon Tack
    @jihoontack
    Jul 27
    Excited to share my first work after joining @MSFTResearch! LLMs have entered the agentic era, and we now collaborate with agents on complex tasks over many turns of interaction. But does your agent actually follow what you intended? We show where agents get lost: LLMs Get Lost
    Image
  • @PhilippeLaban
    Philippe Laban
    @PhilippeLaban
    Apr 21
    New paper! LLMs Corrupt Your Documents When You Delegate LLMs are enabling a new way of working: delegated work, where users supervise an LLM as it edits documents on their behalf. Delegation requires trust: does the LLM complete tasks without introducing errors? We simulate
    Image
    00:00
    43
  • @PhilippeLaban
    Philippe Laban
    @PhilippeLaban
    Feb 24
    LLMs *Still* Get Lost In Multi-Turn Conversation. We re-ran experiments with newer models. Performance still drops, but with modest gains: mostly from improvements on the Python coding task. Also: Lost in Conversation will be presented at ICLR 2026 🎉🇧🇷
    Image
    15
  • @PhilippeLaban
    Philippe Laban
    @PhilippeLaban
    Feb 10
    Very cool work!
    @allen_ai
    Ai2
    @allen_ai
    Feb 10
    LLMs often generate step-by-step instructions, from real-world tasks (how do I file taxes?) to plans for AI agents. Improving this is hard: outputs can sound fluent for steps that don't work, and current datasets cover few domains. How2Everything evals/trains for this at scale.
    Image
Advertisement
Advertisement