C2Dex Turns Phone Videos Into Dexterous Hand Trajectories
Collecting clean demos for multi-finger hands is expensive. Phone video of people manipulating objects is not โ if you can turn noisy monocular hand-object footage into something a robot hand can actually run.
C2Dex, submitted to arXiv on Aug 7, 2026 (arXiv:2608.07045) and aimed at IEEE RA-L, is a video-to-dexterous-manipulation pipeline from Nanjing University and collaborators. The core trick is not better per-frame pose estimation. It is stable object-side contacts: aggregate noisy frame-wise contact observations into the canonical object frame, then use those contacts both to clean the human trajectory and to retarget it onto a robot hand.
What they built
C2Dex has two coupled modules:
- Contact-consistent HOI reconstruction โ start from Dyn-HaMR hand estimates, SAM 3D object geometry, and ProxyPose object poses; filter contacts with silhouette overlap and normal consistency; cluster contacts across stable temporal segments with DBSCAN; refine MANO trajectories with contact, SDF penetration, and regularization losses.
- Contact-interaction-preserving retargeting โ transfer stable contacts to the target hand (Inspire in the paper), preserve local hand-object geometry with Laplacian interaction optimization, then refine in Isaac Gym with residual RL (ManipTrans-style).
The project page (k-jie.github.io/C2Dex) shows open-loop replay on a Unitree G1 with an Inspire hand on eight contact-rich tasks: hang cap, pour juice, sweep litter, wipe board, stack cups, and more. Twenty-four human demos feed the pipeline with no per-task policy training and no teleop on the robot.
Numbers that matter
Under identical evaluation criteria on monocular demos:
| Dataset | C2Dex (relaxed) | Strongest baseline |
|---|---|---|
| DexYCB (45 seq) | 57.78% (26/45) | 17.78% (Do As I Do) |
| TACO (30 seq) | 26.67% (8/30) | 10.00% (Do As I Do) |
Strict rotation thresholds still leave C2Dex far ahead (55.56% / 23.33%). On retargeting alone with ground-truth human HOI input, max penetration on DexYCB drops from 22.92 mm (DexPilot) to 3.99 mm. Ablations show removing cross-frame contact consistency collapses DexYCB success back to 17.78%.
A Humanโs Take
Iโm so here for the boring object-space contact bank. Most video-to-dexterity stacks fail because contacts jitter every frame and retargeting matches fingertips instead of where the object actually gets touched. If C2Dex-style stable contacts hold up outside author-recorded demos and clean object meshes, phone footage stops being a research novelty and starts looking like a data factory for hands.