Pinned
If you want a vision encoder for dexterous manipulation, what should be the most important part to model? ๐ค
Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for






