|
Antonios Tragoudaras
Mail |
LinkedIn |
Google Scholar |
GitHub |
X |
CV
Last updated: August 2026
ABOUT ME
I am a PhD candidate at the Astra-Vision lab,
Inria Paris Centre, where I work on
video foundation models for physical world understanding.
My research interests lie in making robots and intelligent systems understand and model
the real world we live in. I focus on developing AI systems that can perceive, reason about, and
interact with physical environments through a deep understanding of the underlying physical
principles and dynamics.
|
|
RESEARCH PROJECTS
Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems
Antonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonios Vozikis, Thijmen Nijdam, Derck W. E. Prinzhorn, Mark Bodracska, Nicu Sebe, Andrii Zadaianchuk, Efstratios Gavves
International Conference on Machine Learning (ICML), 2026
Links:
Paper
| Code
| Dataset
| Website
We introduce Morpheus, a physics-informed evaluation framework for assessing whether video generative models understand Newtonian dynamics. Morpheus uses 130 real-world videos of physical phenomena and interpretable metrics grounded in conservation laws to measure physical plausibility, showing that todayโs models struggle to encode physical principles even when they generate visually convincing videos.
|
|
Physics-Informed Representation Alignment
Antonios Tragoudaras
OpenReview (pre-print), 2025
Links:
Paper
| Website
We introduce PIRA (Physics-Informed Representation Alignment), a method for grounding physics into video diffusion models by distilling explicit physical proxy signals โ optical flow, depth, segmentation masks, and gravity maps โ through the modelโs own native 3D-VAE encoder. This yields teacher representations that are inherently compatible with the modelโs noisy latents, avoiding the latent-space mismatch of external-teacher distillation. Alignment is performed relationally, matching pairwise token-similarity matrices between student and teacher, and demonstrated on free-fall dynamics with CogVideoX. PIRA is the foundation of an active research thread scaling grounding to multiple interacting objects and richer physical cues.
|
|
|