Check out Jihoon's very interesting work on simulating long "situated" conversations where the user can change their mind, switch tasks, underspecify, etc.
We need more work evaluating model performance outside of the single-turn, single-task, lab-like setting.
Excited to share my first work after joining @MSFTResearch!
LLMs have entered the agentic era, and we now collaborate with agents on complex tasks over many turns of interaction.
But does your agent actually follow what you intended?
We show where agents get lost: LLMs Get Lost




