Yuan Xu*1 · Youheng Shi*1 · Chengyang Li2,3 · Wentao Zhu2 · Yizhou Wang1
1Peking University 2Eastern Institute of Technology, Ningbo 3Shanghai Jiao Tong University
* Equal contribution
TL;DR: VLA fine-tuning on limited robot data erodes its inherited semantic structure and undermines generalization. Inspired by mirror neuron theory, we probe this erosion and reveal its correlation with model performance. We introduce a plug-and-play semantic alignment method that consistently improves performance on simulation benchmarks and real-robot tasks.
![]() |
![]() |
![]() |
![]() |
| pick up the grapes and place it on the plate |
stack the cups | place the toy bear into the box and close both side flaps |
place the sponge in the cabinet and close the door |
![]() |
![]() |
![]() |
![]() |
![]() |
| Language Variation | Novel Object | Visual Distraction | Position Variation | Compositional Tasks |
| Category | Supported |
|---|---|
| VLA Models | pi0, SpatialVLA |
| Simulators | LIBERO, SimplerEnv |
| Datasets | BridgeData V2 (RLDS), LIBERO (LeRobot) |
@misc{xu2026semanticanchoringroboticaction,
title={Semantic Anchoring for Robotic Action Representations},
author={Yuan Xu and Youheng Shi and Chengyang Li and Wentao Zhu and Yizhou Wang},
year={2026},
eprint={2607.13597},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2607.13597},
}This project builds on openpi, SpatialVLA, and StarVLA. We sincerely appreciate the work their authors have shared with the community.









