Spirit AI Puts Moz1 on CATL Lines and Bets on Messy Human Data
Spirit AI’s pitch this week is not another backflip. Co-founder and chief scientist Gao Yang told Reuters that the weak link is still the brain, and that the company is paying people to generate the ugly motion data models actually need.
A date, a robot, a thousand suits
Humanoids Daily and Finimize both reconstruct the 18 September Reuters interview from Beijing. Gao, also an assistant professor at Tsinghua, said:
“We anticipate reaching the GPT-3.0 milestone by mid-2027. You will be able to speak to a robot in natural language, and it will execute a series of reasonable physical actions to attempt the task.”
That is his analogy, not a public benchmark. He split the timeline: next one to two years for industrial jobs, simpler commercial service after that, homes “far harder.” Reuters, as quoted, puts Spirit robots at 90% success on simple tasks in structured living-room settings, and still stuck on things like unscrewing a bottle cap and on tasks that look new. Trial counts and intervention rules are not in the write-ups.
On the floor, Reuters reports tens of Moz1 wheeled humanoids on production lines at battery maker CATL and retailer JD.com, which is also an investor. Moz1 is the industrial body; Moz2 is the commercial-service follow, per Spirit’s own product sequence as summarized by Humanoids Daily.
Data: about 1,000 contractors nationwide wear capture gear in homes and factories. Gao argues simulators handle rigid bodies and choke on flexible stuff like cables, so Spirit leans on real recordings. The counterintuitive claim: “dirty data,” varied imperfect motion, taught models faster than only clean repeats.
The company is about 300 people, founded in 2024. Reuters, via Finimize and HD, puts funding above $670 million and valuation at 20 billion yuan (~$2.9 billion). Gao declined to talk IPO. Spirit has also posted Spirit-v1.5 code and checkpoints on GitHub, which is a concrete artifact next to the forecast.
A Human’s Take
I will remember the thousand contractors and the Moz1s at CATL before I remember “GPT-3 for robots.” Mid-2027 is now a calendar invite. The test is an unfamiliar spoken request, on a line that was not in the capture set, without a human hovering. Bottle caps first.