DEX-X
Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction
Anonymous Authors
DEX-X system overview

DEX-X learns deployable visual-tactile dexterous manipulation policies from human video demonstrations. It transfers human hand-object interactions into simulation, where tactile-aware reinforcement learning enriches visual demonstrations with contact information and learns deployable manipulation policies. The resulting policies transfer zero-shot to real-world grasping and tool-use tasks.

TL;DR
Human videos show what dexterous manipulation looks like, but not what it feels like. DEX-X replays monocular human demonstrations in simulation to recover the missing contact signal, and distills the result into visual-tactile policies that transfer zero-shot to a real hand-arm platform — with no robot-side data collection.
Abstract
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection?

We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing.

We demonstrate zero-shot sim-to-real transfer on dexterous hand-arm platforms across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed across the evaluated tasks.

Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.
Pipeline
DEX-X pipeline

Training framework of DEX-X. After transferring human demonstrations into simulation, we train a privileged state-based expert using RL with demonstration references, object states, and tactile contact information. The expert is then distilled into a multi-task visual-tactile policy operating on a unified contact point cloud representation, which embeds fingertip tactile feedback into the scene geometry and fuses vision, touch, and proprioception for policy learning. The resulting policy transfers zero-shot to diverse real-world dexterous manipulation tasks.

Hardware Setting
Real-world hardware setup: Franka Research 3 arm, Sharpa HA4 hand and RealSense camera

Real-world hardware setup. We evaluate the proposed framework on a Franka Research 3 (FR3) arm with a Sharpa HA4 dexterous hand (29 DoF per arm; 58 DoF in the bimanual configuration). Visual observations are provided by a fixed RealSense depth camera, while tactile feedback is collected from fingertip tactile sensors. All policies are trained in IsaacLab and deployed at 30 Hz.

Simulation Rollouts
Squeegee — collect sand
Rotate squeegee
Use hammer
Pick up cube
Pour from cup
Peg insertion
Bimanual Simulation Rollouts
Uncap jar
Cup handover
Real-world Deployment
The distilled visual-tactile policy is deployed zero-shot on the real hand-arm platform across grasping and contact-rich tool-use tasks. Scroll the gallery below, or use the arrows, to browse all real-world rollouts.
Zero-shot to Unseen Objects
The distilled visual-tactile policy is deployed without any fine-tuning on objects that were never seen during training, spanning novel shapes, colors, and material appearances.
Blue duck
Grey cube
Yellow duck
Square cube