Noitom Drops 617.5 Hours of High-Precision Human Motion for Humanoids
Noitom Robotics released HiPHI at the World Robot Conference in Beijing: 617.5 hours of high-precision optical motion capture from 132 performers, captured at 90 Hz with sub-millimeter marker tracking. The dataset is on Hugging Face now, free for research under the ModalityNet Open Research License.
The company calls it the first public drop of its World Compiler approach: make the physical world learnable for machines. The paper is arXiv:2608.16222.
What is in the box
GlobeNewswire and the project page agree on the split:
- 617.5 hours released, including left-right mirrored counterparts
- 371.8 hours of whole-body human motion
- 245.7 hours of human-object interaction, with each object’s trajectory and mesh recorded in sync
- 308.7 hours of original capture and 200.1 million frames at 90 Hz
- 40 real-world objects across 12 categories, masses from 0.45–6.25 kg
- Organized around FrameNet motion units: 22 frames, 214 Frame–LU labels
Policies trained on HiPHI run on a physical Unitree G1: running, sitting, crawling, carrying a box, and pulling a suitcase, according to the press release.
Why they opened it
“The bottleneck in physical AI is not how much data exists, but how much of it a machine can actually learn from,” said Dr. Tristan Ruoli Dai, Founder and CEO of Noitom Robotics.
Internet video is huge and physically sloppy. Lab MoCap is precise and usually tiny, and companies sit on it. Noitom says HiPHI is among the largest high-precision human-motion sets ever made public. Dr. Lei Han, chief of R&D, said the infrastructure produces more than 100,000 hours a year for partners, and that HiPHI is a faithful sample of that pipeline.
The project page reports 1,620 occupied cells and a 14.1% rare-cell long-tail share, versus 10.7% for the closest baseline they highlight. Tracking error keeps falling as they scale unmirrored training from 3 to 300 hours.
Commercial licensing is through modalitynet.com. Noitom says it will keep launching at RO-MAN 2026 in Fukuoka (Aug. 24–28) and plans an omni-modality interaction corpus later this year, plus SMPL and SOMA formats.
A Human’s Take
I like a company that publishes the thing it usually hoards. 617 hours of optical MoCap with object meshes is a gift to anyone training whole-body policies, and the G1 clips (run, crawl, carry, suitcase) are the receipt that matters.
I still want to see the failure cases, not just the coverage maps. If the long tail is real, the next public drop should be the ugly hours: slips, recoveries, and the motions that make a G1 look confused.