EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
Policies with active stereo gaze match or exceed wrist camera setups in performance.
I'm a final year PhD at Berkeley's AI Research Lab, co-advised by Angjoo Kanazawa and Ken Goldberg. I work on robot perception for manipulation, from 3D scene understanding to, more recently, active vision: teaching robots where to look.
I maintain Nerfstudio, a large open-source framework for 3D neural reconstruction, and am supported by the NSF GRFP. I'm currently an intern at Amazon FAR, and previously interned at Google DeepMind on the Gemini Robotics team.
Before Berkeley, I did my undergrad at CMU where I worked with Howie Choset on multi-robot path planning. I also spent time at Berkshire Grey on warehouse automation and NASA JPL, where I was lucky to be involved in early work on a moon rover tech demo.
Policies with active stereo gaze match or exceed wrist camera setups in performance.
A real-to-sim framework for training robot locomotion policies from RGB pixels in IsaacGym using Gaussian splatted scenes.
We train a robot eyeball policy to look around to enable the performance of a BC arm agent. Eye gaze emerges from RL by co-training with the BC agent and rewarding the eye for correct arm predictions.
A self-improving framework that builds intuition for predicting 3D object configurations through iterative observation, prediction, and optimization cycles.
Tracking and manipulating irregularly-shaped, previously unseen objects using a single stereo camera with Gaussian splatting for object pose estimation and language-driven manipulation.
Object-centric visual imitation from a single video by 4D-reconstructing articulated object motion and transferring to a bimanual robot.
Hierarchical grouping in 3D by training a scale-conditioned affinity field from multi-level masks.
LERF's multi-scale semantics enables 0-shot language-conditioned part grasping for a wide variety of objects.
A modular PyTorch framework for Neural Radiance Fields research with plug-and-play components and real-time visualization tools.
A Python library for interactive 3D visualization in the browser, with an imperative API for building scenes and GUIs for robotics and computer vision.
Grounding CLIP vectors volumetrically inside a NeRF allows flexible natural language queries in 3D.
We collect spatially paired vision and tactile inputs with a custom rig to train cross-modal representations. We then show these representations can be used for multiple active and passive perception tasks without fine-tuning.
NeRF functions as a real-time, updateable scene reconstruction for rapidly grasping table-top transparent objects. Geometry regularization speeds and improves scene geometry, and a NeRF-adapted grasping network learns to ignore floaters.
Fluorescent paint enables inexpensive (<$300) and self-supervised data collection of dense image annotations without altering objects' appearance.
A sliding-pinching dual-mode gripper enables untangling charging cables with manipulation primitives to simplify perception coupled with learned perception modules.
Using NeRF to reconstruct scenes with offline calibrated camera poses can produce graspable geometry even on reflective and transparent objects.
Combining active perception with behavior cloning can reliably hand a surgical needle back and forth between grippers.
A framework for multi-agent pathfinding that combines reinforcement and imitation learning to teach fully decentralized policies, validated with up to 1024 agents.
The result of lots of COVID boredom, I designed, built, and programmed these from scratch (before vibe coding was a thing!).
Miniature self-balancing robot that follows waypoints using model-predictive control for balance and pure pursuit for path following.
Indoor mapping robot with a planar lidar that autonomously mapped my house using a Google Cartographer-inspired SLAM algorithm, alongside a grid path planner for exploration and DWA controller for path following, all on-board a Raspberry Pi.