EGR Stops a VLA From Listening to the Wrong Camera
A VLA with three cameras will happily bind a walk to the floor texture in a wrist view. Yue Yang (UNC Chapel Hill / MERL) and colleagues call that modality entanglement. Their paper, posted 2 September, adds a training loss that asks, per frame, which sensor actually sees the task.
They name the loss Evidence-Gated Regularization (EGR). It does not change the architecture and adds zero cost at inference. The backbone is π0.5.
What EGR actually does
For each frame and each sensor, EGR builds an evidence score. Cameras get it from how much of the focal object is visible. GelSight gets it from contact. That score gates two consistency terms:
- Invariance on low-evidence sensors: corrupt them and the action should not move.
- Sufficiency on high-evidence sensors: keep only that sensor and the action should still match the full observation.
Random modality dropout is the obvious baseline. On their BEHAVIOR-1K skill set it does not fix the entanglement. EGR does, at least relative to vanilla π0.5.
Simulation, 47 rollout skills (11 nav + 36 manipulation), 20 rollouts each:
| Condition | vanilla π0.5 | EGR |
|---|---|---|
| Manipulation, all sensors | 12.5% | 16.4% (+31%) |
| Corrupt the useless camera | 9.4% | 16.5% (+75%) |
| Keep only the useful camera | 2.8% | 6.1% (+120%) |
| Nav, single useful sensor | 15.5% | 37.3% (+141%) |
Clean nav stays about flat (42.3% → 40.9%). The gain is under corruption, which is the point.
Two real robots
They then run the same loss on hardware that does not look alike.
Bi-manual Kinova, three RGB cameras, four tasks × four conditions × 10 trials. Averaged across tasks, EGR takes RealDistractor success from 30% to 85% (+183%). UsefulOnly (only the informative cameras survive) goes 25% → 72.5%. Clean goes 70% → 80%.
MELFA ASSISTA with one RGB camera and two GelSight sensors, two inspection-style tasks (board crack, pencil serial code). SingleUseful goes 40% → 90%. RealDistractor goes 55% → 70%. Clean is a wash within trial noise (90% → 85%).
The authors are honest about the evidence function: it needs task structure you can name (visible objects, a crack signature). In-hand reorientation where the cue itself must be learned is future work.
A Human’s Take
I have watched too many wrist cameras train a policy to “see” a table leg that is not the job. EGR is a training-time scolding: if this camera is looking at the floor, stop using it. The Kinova distractor jump is the number I care about, because that is a physical object the policy never met. The sim success rates are still low in absolute terms. This is a robustness patch, not a new manipulator. Use it that way.