ECCV 2026 Workshop · Malmö

World Models in the Lp:
Towards Application-Driven World Model Evaluation

Tuesday, September 8 — Half-Day Workshop (pm)

Room: Malmö Arena Terrassen

01

Overview

The next challenge for intelligent systems is not only perception but also anticipating and acting in physical space under interventions and uncertainty. World models support simulation, data generation, and planning by rolling out futures conditioned on observations and actions, in applications such as robot manipulation, autonomous driving, and embodied navigation.

Yet recent rapid progress in world models has outpaced evaluation: beyond visual fidelity, we must assess controllability, physical plausibility, and robustness to distribution shift. As world models are getting embedded in larger systems and are effectively used in-the-loop for planning or training and evaluating perception and control policies, open-loop benchmarks provide limited insight. In such settings, errors compound, and missing controllability, physical plausibility, or robustness under distribution shift directly translate into inaccurate rollouts, suboptimal decisions, degraded downstream performance, or potentially catastrophic errors in safety-critical applications. Such properties are task-dependent and, thus, require task-specific evaluation criteria.

02

Topics of Interest

Our workshop rethinks world model evaluation from an application perspective and aims to address key open questions:

01

Key capabilities under interventions

What are key capabilities of world models and their trade-offs? E.g., efficiently predicting plausible futures under interventions, generating conditioned high visual-fidelity videos, or serving as a realistic simulator for closed-loop agent training.

02

Evaluation protocols and benchmarks

How to design evaluation protocols and benchmarks for properties such as controllability, physical plausibility, and robustness? And how can exploration of these insights inform world model training and design?

03

Failure modes in downstream systems

How to identify systematic failure modes that emerge when world models are embedded in downstream systems and how to integrate imperfect world models, especially under challenging scenarios?

03

Call for Papers

We invite submissions of research papers related to the application of world models in other downstream tasks, their benchmarking, and evaluation. This is a Nectar track, i.e., papers will not be published in the ECCV 2026 Workshop Proceedings. The goal is to bring together researchers from different communities in a poster session and to promote papers in this field that were already accepted in or submitted to a previous conference (CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, RSS, CoRL).

Submissions are handled via this Google form. The submission should include a reference to the already published paper or a single PDF containing the final version (not anonymized). Submissions are processed on a rolling basis, with feedback within a week after submission.

All accepted papers will be presented in a poster session. A Best Paper Award will be selected among all accepted papers and announced at the start of the workshop.

Important Dates:

Submission Deadline
September 4, 2026
Acceptance Decision
Rolling basis within 1 week
04

Accepted Papers — Nectar Track

The following papers have been accepted to the Nectar track. Best Paper Award for Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?

🏆 Best Paper Award

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?

Dingrui Wang, Zhihao Liang, Hongyuan Ye, Zhexiao Sun, Zhaowei Lu, Yuchen Zhang, Yuyu Zhao, Yuan Gao, Marvin Seegert, Finn Schäfer, Haotong Qin, Wei Li, Luigi Palmieri, Felix Jahncke, Mattia Piccinini, Johannes Betz

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham, Jeffrey Hawke

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

Shweta Mahajan, Shreya Kadambi, Hoang Le, Rajeev Yasarla, Apratim Bhattacharyya, Munawar Hayat, Fatih Porikli

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

Haozhuang Chi, Daosheng Qiu, Hao Su, Haochen Liu, Zirui Li, Haoruo Zhang, Chen Lv

DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation

Ziyu Shan, Zhenyu Wu, Xiaofeng Wang, Zheng Zhu, Ziwei Wang

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy, Umar Iqbal, Juergen Gall

Lifting Embodied World Models for Planning and Control

Alex N. Wang, Trevor Darrell, Pavel Izmailov, Yutong Bai, Amir Bar

Runway Characters: Real-Time Expressive AI Characters from a Single Image

Rohan Agarwal, Robin Andeer, Alex Armbruster, Bennett Arthur, Michail Doukas, Piero Esposito, Qiming Fang, Aron Filkey, Bryan Fox, Anastasis Germanidis, Jonathan Granskog, Corina Gurau, Robin Kahlow, Taras Khakhulin, Nasir Mohammad Khalid, Tomasz Łakomy, Mike Leisz, Kathleen Lewis, Victor Luo, Daniil Merkulov, Nicolas Neubert, Ryan Phillips, Yining Shi, Axel Örn Sigurðsson, Erica Simmons, Mark Sturley, Ziyi Tang, Michail Tarasiou, Jamie Umpherson, Cindy Wang, Jimei Yang, Hudson Yeo, Wei Zhang, Zejia Zheng

05

Schedule

13:30 – 13:40 Opening Remarks & Best Paper Award Announcement
13:40 – 14:10 Invited Talk 1. Ranjay Krishna: Visual reasoning by simulating future steps.
14:10 – 14:40 Invited Talk 2. Vincent Sitzmann: Video models for embodied intelligence.
14:40 – 15:10 Invited Talk 3. Laura Leal-Taixé: Multi-modality and controllability in generative simulation.
15:10 – 16:10 Coffee & Poster Session (Boards 88–97)
16:10 – 16:40 Invited Talk 4. Carl Doersch: Towards better motion modeling for robots and animals.
16:40 – 17:10 Invited Talk 5. Amir Bar: Planning with world models.
17:10 – 17:50 Panel Discussion: Amir Bar, Carl Doersch, Vincent Sitzmann, Andrea Tagliasacchi. Moderator: Thu Nguyen-Phuoc.
17:50 – 18:00 Ending Remarks
06

Invited Speakers

Ranjay Krishna

University of Washington, Microsoft AI

Laura Leal-Taixé

NVIDIA, University of Toronto

Carl Doersch

Google DeepMind

Amir Bar

Imperial College London, AMI Labs

Andrea Tagliasacchi

Simon Fraser University, Wayve

07

Organizers

Olaf Dünkel

MPI-INF

Nhi Pham

MPI-INF

Ken-Joel Simmoteit

TU Darmstadt

Thu Nguyen-Phuoc

Prometheus

Adam Kortylewski

CISPA Helmholtz Center

08

Advisory Board

Jan Peters

TU Darmstadt, DFKI

Alan Yuille

Johns Hopkins University