GIFT Trains Robot Policies to Keep Geometry, Affordances, and Goals
A team from the Chinese Academy of Sciences, Tsinghua, Fudan, and the National University of Singapore posted GIFT on 3 September (arXiv:2609.04193): a training recipe that forces a robot policy’s intermediate visual features to keep three control-relevant structures. Geometry for “can I move there,” affordance for “what to grab and how,” and goals for “which pixels actually matter for the instruction.”
The authors call the usual mismatch the action-sufficiency gap. Vision-language pretraining and world models give rich pictures. Those pictures still dump a lot of stuff a gripper does not need.
What they actually train
GIFT does not invent a new action head. It wraps three existing policy families and keeps each family’s action math:
- GIFT-VLA: a semantics-centered vision-language-action policy
- GIFT-WAM-Fast: a world-action model that dumps a current-frame action
- GIFT-WAM-IDM: an inverse-dynamics world-action model that imagines a trajectory first
The extra losses sit on intermediate tokens. Geometry is aligned to a frozen VGGT teacher. Affordance heads predict entity roles, object-centric end-effector poses, and gripper closure. A goal head reconstructs instruction-relevant image regions. At deployment the extra heads can stay off. The paper’s default is no-injection: the auxiliaries supervise features, they do not get concatenated into the action path.
On zero-shot LIBERO-Plus (seven distribution shifts, no fine-tune), the project page reports:
- GIFT-VLA 79.6% vs StarVLA-OFT 75.0% (+4.6 points)
- GIFT-WAM-Fast 72.6% vs Fast-WAM 60.0% (+12.6)
- GIFT-WAM-IDM 87.8% vs Fast-WAM-IDM 82.6% (+5.2)
On RoboCasa, the same three land at 61.4%, 83.6%, and 82.3%. Articulated-object tasks are where the world-action variants move the most: +21.3 and +24.6 points over their Fast-WAM counterparts.
Hardware, not just sim
The real suite is four tasks, ten trials each, on the two platforms above: color-conditioned placement, layer-conditioned placement, cabinet work, and bimanual test-tube insertion. GIFT-WAM-IDM hits 87.5% in the original settings and 67.5% across perturbation conditions, against 52.5% and 15.0% for Fast-WAM-IDM (40 original trials, 40 perturbed). Level-2 perturbations include a rotating colored lamp, a swapped tabletop, and a 15–20° rack rotation.
Ablations on the project page say each single guidance signal helps, all three together win, and the no-injection variants beat the injection ones on both LIBERO-Plus and RoboCasa.
A Human’s Take
I like that they left the action decoder alone. If your VLA already works, the useful question is whether the features it sees are the ones that decide a grasp, not whether you can stack another transformer on top. The real-arm gap under perturbation is the number I would take into a lab meeting. Eighty-seven percent on a clean table is nice. Sixty-seven percent when someone spins a lamp and rotates the rack is closer to a shift.