AI

DreamX-Phi Predicts What a Bimanual Arm Will See Next

Shar Hendrix 4 min read

The DreamX Team posted DreamX-Phi 1.0 on August 13: a video world model that takes one RGB frame, a language instruction, and a prescribed bimanual action sequence, then predicts the next observations. On a pinned WorldArena 2.0 snapshot dated August 12, their submission ranked first on Track 1 (EWMScore-P 60.65) and tied for second on Track 2 (67.19% Adjust Bottle success).

The paper and the public GitHub repo (AMAP-ML/DreamX-Phi) match on those ranks. Weights stay closed until the WorldArena 2.0 IROS challenge ends.

DreamX-Phi 1.0 overview of action-conditioned video prediction
Paper teaser: one frame plus left/right SE(3) paths in, predicted rollout out. Source: DreamX Team, arXiv:2608.13489.

Pretty video is not enough

A convincing rollout can still move the wrong arm or drop the object. DreamX-Phi is built on Wan2.2-TI2V-5B. The authors inject per-arm SE(3) transforms into attention with PRoPE-style geometric encoding so each arm keeps its identity. Gripper opening is a separate scalar bias, because it is not a rigid transform.

They add three extra checks during training:

  • a depth branch (Depth Anything 3 targets) so scene geometry does not drift
  • SAM3 object masks that reweight the RGB loss onto the thing being grasped
  • a frozen V-JEPA teacher that keeps object relations consistent over time

DMD (distribution-matching distillation) then compresses the multi-step generator into a few-step student.

Predicted WorldArena clean-scene rollouts of bimanual robot arms
Qualitative Track 1 rollouts on clean RoboTwin 2.0 scenes. Source: DreamX Team, arXiv:2608.13489.

What they trained on

The corpus mixes Ego4D (3,700 h), AgiBot World 2026 (1,900 h), InternData-A1, Cosmos3-DROID, RoboCOIN, and 25,000 RoboTwin 2.0 clips. They drop mobile-base and parked segments and keep failed executions on purpose. After filtering, the AgiBot imitation split is 178.7 hours.

Track 2 uses the world model as a rollout environment to train a π0.5 policy, then tests that policy on held-out Adjust Bottle episodes. The paper is clear this is not DreamX-Phi acting as a closed-loop controller.

DreamX-Phi training diagram with PRoPE, depth, SAM3, and DMD
Training stack: geometry-aware attention, depth/object losses, then few-step distillation. Source: DreamX Team, arXiv:2608.13489.

A Human’s Take

I care that they treated “the wrong arm moved” as a failure mode, not a footnote. Leaderboard snapshots expire; the useful bit is the interface: keep each arm’s rigid path in the attention, then supervise the object so it does not teleport. I will believe it more when the weights ship and someone runs it on a real bimanual bench that is not RoboTwin.

Sources