DEX-X
Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction
Accepted at CoRL 2026
Ruoqu Chen, Feixiang Ruan, Liu Cao, Zihao Wang, Botian Xu, Shiqin Tong, Jiajun Liu, Mingzhi Pei, Chenyu Zhang, Wanli Xing, Kaifeng Zhang, Mengdi Xu
DEX-X system overview

DEX-X learns deployable visual-tactile dexterous manipulation policies from human video demonstrations. It transfers human hand-object interactions into simulation, where tactile-aware reinforcement learning enriches visual demonstrations with contact information and learns deployable manipulation policies. The resulting policies transfer zero-shot to real-world grasping and tool-use tasks.

TL;DR
Human videos show what dexterous manipulation looks like, but not what it feels like. DEX-X replays monocular human demonstrations in simulation to recover the missing contact signal, and distills the result into visual-tactile policies that transfer zero-shot to a real hand-arm platform — with no robot-side data collection.
Abstract
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection?

We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing.

We demonstrate zero-shot sim-to-real transfer on dexterous hand-arm platforms across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed across the evaluated tasks.

Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.
Pipeline
DEX-X pipeline

Training framework of DEX-X. After transferring human demonstrations into simulation, we train a privileged state-based expert using RL with demonstration references, object states, and tactile contact information. The expert is then distilled into a multi-task visual-tactile policy operating on a unified contact point cloud representation, which embeds fingertip tactile feedback into the scene geometry and fuses vision, touch, and proprioception for policy learning. The resulting policy transfers zero-shot to diverse real-world dexterous manipulation tasks.

Hardware Setting
Real-world hardware setup: Franka Research 3 arm, Sharpa Wave hand and RealSense camera

Real-world hardware setup. We evaluate the proposed framework on a Franka Research 3 (FR3) arm with a Sharpa Wave dexterous hand (29 DoF per arm; 58 DoF in the bimanual configuration). Visual observations are provided by a fixed RealSense depth camera, while tactile feedback is collected from fingertip tactile sensors. All policies are trained in IsaacLab and deployed at 30 Hz.

Simulation Rollouts
Squeegee — collect sand
Rotate squeegee
Use hammer
Pick up cube
Pour from cup
Peg insertion
Bimanual Simulation Rollouts
Uncap jar
Cup handover
Real-world Deployment
The distilled visual-tactile policy is deployed zero-shot on the real hand-arm platform across grasping and contact-rich tool-use tasks. Scroll the gallery below, or use the arrows, to browse all real-world rollouts.
Zero-shot to Unseen Objects
The distilled visual-tactile policy is deployed without any fine-tuning on objects that were never seen during training, spanning novel shapes, colors, and material appearances.
Blue duck
Grey cube
Yellow duck
Square cube