Bernstein says Figure AI's Helix 2.5 demonstrates 'scene generalization,' not 'task generalization,' with success rates still far from commercial viability

Stock News
Yesterday

Bernstein's global automation team has released a fresh analysis centered on Figure AI's Helix 2.5 launch this week, offering a measured yet decisive take on the latest strides in robotic 'brain models.' The firm concludes that while Helix 2.5 proves a humanoid's 'brain' can acquire general skills from human data, what it has conquered is 'scenes,' not 'tasks' — the robot still depends on task-specific training to function, and a 56% success rate remains a long way off from being commercially useful.

Across 30 unfamiliar homes and 420 trials, Helix 2.5 achieved a 56% success rate. According to Figure's official release on September 17 and the Bernstein report, Helix 2.5 executed three long-horizon, whole-body tasks — tidying a living room, folding towels, and making a bed — in 30 homes in the San Francisco Bay Area that it had never encountered before. Throughout the process, Figure collected no data in these environments, performed no fine-tuning, and made no adaptations. Prior robot demonstrations typically took place in trained, familiar settings. For instance, Helix 02 in February showcased autonomous completion of a multi-minute dishwasher unloading task and ran a logistics operation autonomously for 200 hours, yet those systems relied on data gathered at the robot's actual operating sites. Helix 2.5 aims to answer a harder question: can a humanoid robot, using full-body autonomy, start working immediately in a never-seen home? The result is a report card showing progress alongside clear gaps. Across the 30 households, 420 trials yielded 237 successes for an overall success rate of 56%, using an all-or-nothing criterion — a task counts only if completed in full. Broken down by task, Bernstein cites Figure's data: towel folding hit 62%, bed-making 67%, and living room tidying just 40%. Living room tidying proved the most difficult, with only a four-in-ten success rate. The report also highlights a promising signal: the robot exhibited self-correction in unfamiliar settings — stepping back to reposition, adjusting its stance, or moving to the other side of the bed to continue making it. But the all-or-nothing yardstick also means nearly half of all trials ended in failure.

The report attributes Helix 2.5's gains to a shift in data paradigm. Helix 2.5 was pretrained from the start on Index, Figure's proprietary dataset of first-person human task videos. Bernstein stresses that with task-specific data, model architecture, training, and evaluation conditions held constant, the single variable of whether Index pretraining was used lifted zero-shot success from 9% to 56%. This improvement is not simply the result of adding more household examples. Figure also disclosed that no evaluated task accounted for more than 1.90% of Index's pretraining data. At the same time, Helix 2.5 reached comparable performance using only half the task-specific data that Helix 02 required for a certain behavior, then generalized that behavior across 30 unseen homes. Of greater interest is the scaling effect. The report states that Figure has gathered preliminary evidence that as 'human-to-robot' pretraining data expands, action-prediction loss steadily declines — doubling Index pretraining data repeatedly produced smooth improvements in downstream robot action prediction, allowing the loss for the largest run to be forecast to four decimal places before training even began. Figure's website describes this as the first scaling law for 'human-to-humanoid' transfer ever measured on a humanoid platform. Still, the report cautions that this experiment measures action-prediction loss, not real-world task completion rates.

The data engine itself is compounding. According to the Bernstein report, as of August 2026, Index contains 23 million videos, adding roughly 35 minutes of human experience per second — also cited as about 5.75 years of human experience per day, two ways of stating the same figure. Every 1,000 hours of collected data corresponds to 373 tasks, 1,146 manipulated objects, and 116 environments. Among 1 million hours of video, there are 373,000 unique tasks and 1.1 million unique manipulated objects. Weekly active data contributors have reached 44,000, climbing rapidly from fewer than 100 in early May. Index launched on August 25 and has surpassed 264,000 downloads. Figure has paid out $15 million to data contributors and committed $3.5 billion in compute resources to training the Helix series.

The key divergence: scenes have generalized, tasks have not. Bernstein characterizes this progress as 'scene generalization' rather than 'task generalization.' The report notes that Figure's robot was specifically trained on these three tasks, and the objects and their arrangements across the 30 homes varied only within a limited range. As a result, true zero-shot task generalization 'still has a long way to go.' The report links Helix 2.5 to the company's earlier note titled 'The Emperor's New Brain and the AlphaGo Moment,' arguing it supports three conclusions. First, model advances have dramatically broadened the scope of training data, with first-person human video becoming increasingly usable in pretraining and growing evidence of scaling laws — more is better. Second, real robot data still holds value, but teleoperation is likely to be phased out due to its lack of scalability. Third, the integration of brain and data is deepening — Helix 02's approach of directly inheriting a fixed-weight vision language model is being replaced by models trained from scratch on proprietary data. Other industry developments are viewed as reinforcing this same trend, including AgiBot's recent world action model GE-ACT 2.0, Dyna 2.0 trained on 1 million hours of human video, and Physical Intelligence's π0.7, which demonstrated compositional task generalization — the emergence of novel, untrained skills from combinations of related capabilities.

The next test: from 56% to 99%. The report lays out two specific expectations for Figure AI. First, evidence of 'transferable skills' — after training on one task, the robot can complete similar tasks without additional training. Second, a true 'data flywheel' — automatic performance gains in both training and deployment phases. Bernstein is blunt that a 56% overall success rate is 'far from sufficient for commercialization.' The adage 'practice makes perfect' has yet to hold in robotics: pretraining typically brings success rates to 50%-60%, and finding an efficient path from that level to 99% will be the next major test. The report's conclusion centers on industry pace rather than timing. Progress from Figure, AgiBot, Physical Intelligence, and Dyna sends a consistent message: large-scale deployment beyond general tasks and a few applications will not arrive soon, but humanoid technology — especially brain models — is advancing at a perceptible pace every quarter. The report sums up the current moment in one line: 'the worst of times, and the best of times.'

On investment implications, this global automation report assigns ratings as follows: Keyence, Cognex, Fanuc, Harmonic Drive Systems, and Inovance Technology all receive outperform ratings. The Bernstein market table, based on closing prices on September 18, 2026, and target prices from the report, reflects these positions.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10