On September 17, Figure AI released its next-generation model, Helix 2.5, and conducted a zero-shot test in 30 real Bay Area homes. It succeeded in 237 of 420 trials, for a 56% zero-shot success rate. The evaluation was blind, and no single evaluation task accounted for more than 1.90% of the pretraining data. As a reference, a policy using the same hardware and the same task data—but without Index pretraining—achieved only a 9% success rate. As pretraining data increases, robot performance shows a predictable scaling relationship similar to that of large language models: with other conditions unchanged, pretraining raised zero-shot task success from 9% to 56%, an overall success rate improvement of 522%. This matters because over the past few years, almost every “impressive” home robot demo has relied heavily on site-specific data or human teleoperation. Figure’s previous-generation model could complete long-horizon tasks spanning dozens of consecutive steps, but it still needed to collect data at the target site first. Helix 2.5 seeks to answer a different question: can a robot walk into a home it has never seen before and start working?
文章图片 2
Supporting this shift is the data collection system Index. According to official disclosures, Index has uploaded more than 16 million videos cumulatively, with upload speeds at one point reaching 30 minutes of human work per second. Figure has already paid $15 million for this and plans to invest more than $1 billion over the next 12 months in data and computing power, expanding collection scale by another 100x. Company leadership has stated clearly that to bring humanoid robots into homes, the biggest bottleneck right now is data and computing power.
文章图片 4
When embodied AI bets its scale on data and computing power, the efficiency of the underlying infrastructure—GPU computing platforms, AI chips, and heterogeneous clusters—determines iteration speed. StarWar Cloud’s GPU computing platform and heterogeneous cluster solutions help teams pool distributed compute, schedule it on demand, and make large-scale training and inference investment more controllable, while also making it easier to scale smoothly as data volume grows. Of course, a 56% success rate does not mean the robot will always fail, because the model already has self-correction and retry capabilities in long-horizon tasks. But the evaluation currently covers only three task categories. Expanding the sample size and covering more real-world scenarios remains an essential path from “demoable” to “deployable.” From demo to deployment, embodied AI still has to cross the chaos of real life: blocks on the carpet, charging cables in couch seams, and toys a child has temporarily piled up can all alter a previously viable route. The more models depend on data and computing power, the closer to true scale goes whoever handles those two things more efficiently.