Train a Humanoid from AI Video, Not MoCap: NCKU’s Synthetic Demo Pipeline
Collecting real humanoid demonstrations is expensive, narrow, and weirdly personal—every person does the same task differently. A National Cheng Kung University (NCKU) team proposes skipping MoCap entirely: generate diverse human motion videos from text prompts, retarget them to a Unitree G1, and train RL motion-tracking policies in simulation.
The paper, Learning Diverse Humanoid Tasks via Synthetic Video Scenarios without Real World Data (arXiv:2607.21648, 22 July 2026), is pure sim for now—but the pipeline is concrete.
How the pipeline works
Authors Yun-Hao Tsai, Cong-Thanh Vu, and Yen-Chen Liu describe three stages:
- Prompt-driven video generation — structured instruction prompts plus user task prompts drive Google Veo 3 (and related generative models) to produce full-body, camera-stable human motion clips. Multiple videos per task capture different execution styles.
- Video → humanoid reference — SMPL-X style pose extraction, then retargeting to the humanoid (GMR) with limb scaling, foot-slide fixes, and IK under joint limits.
- Motion stitching + RL tracking — root alignment and joint smoothing to chain clips; DeepMimic-style RL (PPO in NVIDIA Isaac Lab) with imitation rewards on pose, velocity, end-effectors, and CoM, plus limit/smoothness/self-collision penalties.
They report training with 4096 parallel agents on Isaac Lab hardware (Xeon W5-3435X + RTX 4000 Ada in their setup).
What they evaluated
- Motion generation: 50 everyday task prompts × 10 videos each; most everyday motions looked physically plausible; high-speed actions (e.g. backflips) still showed temporal glitches.
- Policy tasks: lie-and-stand, boxing, pick-and-place (including with a 0.5 kg payload). Joint-position MAE roughly 0.04–0.07 m on dynamic tasks; loads hit the upper body harder while lower-body tracking stayed steadier.
The conclusion is careful: the method works in simulation for diverse whole-body tasks; sim-to-real and physical consistency of generated video are left for future work.
A Human’s Take
If generative video can flood the imitation dataset with style variation, humanoid learning stops waiting for a motion-capture studio booking. That is a real data bottleneck unlock—on paper.
I care about two follow-ups: (1) does this survive domain gap to a real G1 without collapsing, and (2) how often Veo invents physics that the retargeter happily teaches the robot. Diversity is worthless if half the demos are soft-body fiction.