Robotics

ADEPT Learns Dexterity Once, Then Inserts Pegs From Cameras

Robb Harlan 5 min read

Most dexterous RL papers relearn “reach, grasp, lift” for every new task. ADEPT tries to pay that bill once.

The paper, posted August 19 as arXiv:2608.19182, comes from NVIDIA and the University of Michigan Robotics group. Authors include Jayjun Lee, Jessica Yin, Ankur Handa, and Nima Fazeli. The project page walks the same pipeline: pretrain on generic object reposing, post-train a specialist, distill a vision (or visuo-tactile) student, deploy zero-shot.

Kuka-Allegro placing a plate in a dish rack and Flexiv-Sharpa inserting a peg, shown as sequential frames
Real-world ADEPT students: dish-rack on a 23-DoF Kuka-Allegro, peg insert on a 29-DoF Flexiv-Sharpa. Source: ADEPT paper / project page.

Pretrain the boring skills

Pretraining is PPO on a reposing task: 16 primitive shapes at random scale, reach-grasp-lift-reorient-transport to a sampled pose. The Kuka-Allegro teacher hits 0.73 success on those primitives, 0.76 on the unseen FMB pegs, and 0.77 on 152 VisDex objects. Flexiv-Sharpa sits a bit lower (0.64 / 0.58 / 0.61).

That prior zero-shots the easy part of peg insertion. On the Kuka-Allegro, success stays above 50% through ADR level 35 and drops to 0% at the actual hole (ADR 50). So post-training only has to learn contact-rich insertion, not grasping from scratch.

Naïve PPO fine-tuning wrecks the prior in a few updates. ADEPT’s recipe is: behavior-clone the actor into the new observation space, freeze it while a fresh critic warms up, then run conservative PPO (actor LR 1e-5, clip down to 0.05). Ablations say the low actor learning rate is the piece that stops collapse.

A full joint-space geometric fabric sits between policy and robot. Unlike earlier fabric work that stuffed the hand into a 5-D PCA grasp, ADEPT commands all 23 (Kuka-Allegro) or 29 (Flexiv-Sharpa) joints and still keeps collision and joint-limit guards.

Lab photos of Kuka-Allegro and Flexiv-Sharpa arms with peg boards, plates, and cameras
Hardware: Kuka iiwa7 + Allegro, Flexiv Rizon + Sharpa, two RealSense cameras each. Source: ADEPT paper.

What actually ran on the bench

Students see stereo RGB. The Sharpa student also gets five fingertip TacMap depth images. No pose tracker, no human demos at deployment.

Real-world, 10 trials per condition, cumulative stages:

  • Kuka-Allegro, FMB star peg: 5/10 full inserts
  • Kuka-Allegro, square/round peg: 3/10
  • Flexiv-Sharpa, square/round, vision only: 3/10
  • Flexiv-Sharpa, square/round, visuo-tactile: 8/10
  • Kuka-Allegro, dish rack: 6/10

Touch is the split. The paper says the vision-only student often cannot tell a grasp landed, so it reopens and loops. With tactile maps it grasped and lifted in every trial.

Cycle time is 5–10 s per trial versus 20–70 s for the FMB parallel-jaw pipeline that uses fixtures and multi-stage regrasps. The authors call that a 2×–14× speedup.

A single-stage distillation student transferred at 0/10. The two-stage curriculum (reposing teacher first, then insertion teacher) is what made sim-to-real land.

Four-panel ADEPT pipeline from reposing pretraining through real-world zero-shot deployment
The four-stage recipe: pretrain, post-train, distill, deploy. Source: ADEPT paper.

A Human’s Take

Pretrain on cubes, then insert a two-legged peg from pixels, with no motion-capture crutch. That is the right research shape.

The 3/10 vision-only inserts on the hard peg still say perception is the bottleneck they named. I care more about the 8/10 with fingertip cameras: if you cannot feel the grasp, you will keep opening the hand. Hands that guess are not ready for a shift.

Sources