Human and Robot Wear the Same Exoskeleton to Teach a 20-DoF Hand
Most wearable-exoskeleton demos record the human and then guess how that motion should land on a robot. SEED-UMI, a CoRL 2026 paper from Peking University and Delta Intelligence, skips the guess. The person and the robot wear the same outer shell.
Joint encoders become a shared measurement. Wrist cameras bolted to the exoskeleton see the same mechanism during collection and rollout. The paper (arXiv:2609.11753, submitted September 10, 2026) reports a 70.0% mean success rate after paired fine-tuning across five contact-rich tasks.
One glove, two bodies
The hardware is 20 independently measured joints across five fingers, matched to a fully actuated Wuji hand on a RealMan RX75 arm. A dorsal Intel RealSense T265 tracks the wrist. A ventral fisheye camera covers about 150°. A parallel four-bar linkage parks the magnetic encoders on the back of the hand so palms and sides stay free for contact.
Mapping is two-stage. First the robot wears the glove and babbles through joints to train an encoder-to-command map. Then human contact-rich motions are replayed on the robot; the gap between the two encoder traces is the supervision. Policies (ACT, Diffusion Policy, π₀.₅) train on raw wrist images and encoder states. No simulation, no RL, no hand segmentation, no inpainting.
Five tasks, one number
Each task gets 100 human demonstrations and 20 autonomous rollouts. After paired fine-tuning, mean success is 70.0%, +12.7 points over babbling-only mapping. Teleoperation, trained on robot-side demos, sits at 71.7%.
- Screw driving: align and hold axial pressure
- AirPods case insertion: index and middle on a small object
- Ball basket throwing: grasp-to-release timing
- Table cleaning: tissue, wipe, discard
- Air freshener spray: hold the can and start a downward press (full spray is not required)
On AirPods, the same operator collected 52 successful demos in 30 minutes with SEED-UMI versus 18 with teleoperation, about 3× throughput. That comparison excludes one-time paired-replay overhead.
The current glove is co-designed for one target hand. Other hands still need geometry and mapping changes. Air Freshener Spray counts a started press, not a full spray.
A Human’s Take
I like the stubborn hardware idea more than the leaderboard. If the camera and the encoder see the same metal on both sides, retargeting stops being a research problem and starts being a calibration. Seventy percent is not a factory shift. It is close enough to teleop that I would rather spend the extra collection minutes on more messy objects than on another open-loop mapper.