Overview
The next challenge for intelligent systems is not only perception but also anticipating and acting in physical space under interventions and uncertainty. World models support simulation, data generation, and planning by rolling out futures conditioned on observations and actions, in applications such as robot manipulation, autonomous driving, and embodied navigation.
Yet recent rapid progress in world models has outpaced evaluation: beyond visual fidelity, we must assess controllability, physical plausibility, and robustness to distribution shift. As world models are getting embedded in larger systems and are effectively used in-the-loop for planning or training and evaluating perception and control policies, open-loop benchmarks provide limited insight. In such settings, errors compound, and missing controllability, physical plausibility, or robustness under distribution shift directly translate into inaccurate rollouts, suboptimal decisions, degraded downstream performance, or potentially catastrophic errors in safety-critical applications. Such properties are task-dependent and, thus, require task-specific evaluation criteria.
Topics of Interest
Our workshop rethinks world model evaluation from an application perspective and aims to address key open questions:
Key capabilities under interventions
What are key capabilities of world models and their trade-offs? E.g., efficiently predicting plausible futures under interventions, generating conditioned high visual-fidelity videos, or serving as a realistic simulator for closed-loop agent training.
Evaluation protocols and benchmarks
How to design evaluation protocols and benchmarks for properties such as controllability, physical plausibility, and robustness? And how can exploration of these insights inform world model training and design?
Failure modes in downstream systems
How to identify systematic failure modes that emerge when world models are embedded in downstream systems and how to integrate imperfect world models, especially under challenging scenarios?
Call for Papers
We invite submissions of research papers related to the application of world models in other downstream tasks, their benchmarking, and evaluation. This is a Nectar track, i.e., papers will not be published in the ECCV 2026 Workshop Proceedings. The goal is to bring together researchers from different communities in a poster session and to promote papers in this field that were already accepted in or submitted to a previous conference (CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, RSS, CoRL).
Submissions are handled via this Google form. The submission should include a reference to the already published paper or a single PDF containing the final version (not anonymized). Submissions are processed on a rolling basis, with feedback within a week after submission.
All accepted papers will be presented in a poster session. A Best Paper Award will be selected among all accepted papers and announced at the start of the workshop.
Important Dates:
Accepted Papers — Nectar Track
The following papers have been accepted to the Nectar track. Best Paper Award for Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
Lifting Embodied World Models for Planning and Control
Runway Characters: Real-Time Expressive AI Characters from a Single Image