Why this workshop?
The rapid transition from large language models (LLMs) as single-turn assistants to interactive agents has created an urgent need for new evaluation methodologies. LLMs are increasingly deployed in high-impact settings such as education, counseling, negotiation, research assistance, and software development, where success depends not only on generating a correct response, but on sustaining effective interactions over extended trajectories.
Evaluating interactive agents directly with real users can be slow, expensive, difficult to reproduce, and hard to scale in expert domains. This has led to growing use of automatic evaluation methods, ranging from rubric-based grading to user simulators, where an LLM simulates user behavior to support evaluation, training, and stress testing. However, many issues with these approaches remain: for example, user simulators may fail to preserve latent user states, reflect diverse human attributes, represent realistic goals, or match the interaction style of real users.
In light of these challenges, this workshop will focus on methods for developing more rigorous, scalable, and scientific evaluation methods for interactive agents.
Topics
We invite contributions on topics including, but not limited to:
- Evaluation protocols for multi-turn assistants, tool-using agents, computer-use agents, collaborative agents, and user-facing systems
- Trajectory-level evaluation, including transcripts, tool calls, intermediate states, final task outcomes, latency, cost, and other operational metrics
- Realistic simulation of users, environments, and interaction partners
- Validation of user simulators as proxies for human behavior and as stress tests for deployed agents
- Grader design, including deterministic checks, model-based rubrics, human evaluation, and calibration between them
- Benchmarks for long-horizon interaction, memory, adaptation, error recovery, and reliability across repeated trials
- Learning from interaction data, production failures, user feedback, and human preference signals
- Safety, fairness, privacy, and ethical considerations in evaluating interactive agents and simulated users
Call for papers
We invite submissions on the topics listed above. Early-stage work is welcome.
- Where to submit. Submissions are open now on our OpenReview submission site. The deadline is August 29, 2026, Anywhere on Earth.
- Format and length. Submissions must use the official NeurIPS 2026 style, available as an Overleaf template. Full papers may be up to 9 pages and short papers up to 4 pages, excluding references and appendices.
- Double-blind review. All submissions are reviewed double-blind, so please anonymize your paper: remove author names and affiliations, and avoid identifying information in the text, acknowledgments, and links.
- Non-archival. The workshop is non-archival and accepted papers will not appear in published proceedings. You may submit work that is currently under review at, or has already been accepted to, another venue, and you remain free to publish it elsewhere afterwards.
- Papers accepted to NeurIPS 2026. Papers already accepted to the NeurIPS 2026 main conference will undergo an expedited review that primarily evaluates their relevance to the workshop themes.
- Opinion papers. We welcome opinion papers. The title must state the opinion and follow the format “Opinion: [Your Title]” — for example, “Opinion: Large Language Models Should Not Replace Peer Review in Scientific Publishing”.
- Presentation. Accepted work will primarily be presented as posters, with a select number of papers receiving spotlight talks, as well as a best paper award.
- In-person attendance. This is an in-person workshop, and we expect at least one author of each accepted paper to attend and present in person, barring unexpected circumstances.
Invited speakers
Schedule
Organizers
Yao Dou
Georgia Tech
Siyan Li
Columbia University
Marwa Abdulhai
Princeton University
Nicholas Tomlin
TTIC
Michel Galley
Microsoft Research
Jacob Eisenstein
Google DeepMind
Alan Ritter
Georgia Tech
Wei Xu
Georgia Tech
Questions? Contact the workshop chairs:
douy@gatech.edu · siyan.li@columbia.edu · marwa_abdulhai@berkeley.edu