AI

GORDON Learns Object-Centric Rewards From Action-Free Video

Robb Harlan 4 min read

Sparse rewards make long-horizon manipulation painful. Raw-pixel reward models break when the background, lighting, or robot body changes.

GORDON (arXiv:2608.03753, submitted Aug 4, 2026) is a graph-based, object-centric reward learner that builds dense progress signals from action-free video demos. Authors are Andrea Protopapa, Davide Buoso, Francesca Pistilli, Georgia Chalvatzaki, and Giuseppe Averta. Project page: andreaprotopapa.github.io/graph-reward-learning.

GORDON comparison: raw-image methods distracted by robot motion vs object-centric graph rewards
GORDON vs raw-image reward methods: object graph and subtask-aware progress. Source: arXiv:2608.03753.

How it works

Each scene becomes a graph of detected objects and spatial relations. A graph neural network embeds those graphs into a task-aligned latent space with self-supervised training. An activity-aware weighted pooling step emphasizes task-relevant objects and downweights robot-dominated motion so the latent is not just “arm moving on camera.”

Dense reward is the distance in that latent space from the current state to demonstrated goal configurations. On long tasks, the reward’s temporal profile shows stage-wise object-state transitions, which GORDON uses for automatic subtask discovery without manual segmentation. Segmented demos then train subtask-specific rewards and specialized policies composed sequentially.

GORDON pipeline: full-task reward, automatic subtask discovery, per-subtask RL, sequential executor
Automatic subtask discovery and sequential RL training from one full-task reward profile. Source: arXiv:2608.03753.

Results (from the paper)

  • Seven manipulation tasks on MAGICAL and ManiSkill3
  • Long-horizon average success rate 74.4%
  • Roughly +35 percentage points vs best learned baseline and +25 p.p. vs oracle (paper’s reported averages)

Short-horizon settings also improve when the object-centric reward is used for RL.

A Human’s Take

Object graphs will not save you if detection fails in a greasy cell — but for sim and clean bench work, “reward from video without actions, then split the long task for free” is the right kind of lazy. I want the open project code run on a third-party ManiSkill suite before I trust the +35 p.p. number outside the authors’ plot.

Sources