Close Menu
Humanoid Analytics
  • Companies
  • Company Tracker
  • Deployment Tracker
  • Funding Tracker
  • Deployments
  • Technology
  • Funding
  • Markets
  • Our Services
Highlights

Figure’s Helix 2.5 Brings Humanoid Robotics Closer to Its ChatGPT Moment

September 18, 2026

Agility’s Digit 5 Adds Scale Features, but Operating Proof Is Still Ahead

September 16, 2026

China Raises the Reported IPO Bar for Humanoid Robotics

September 12, 2026
X (Twitter) Mastodon LinkedIn
Humanoid Analytics
  • Companies
  • Trackers
    • Company Tracker
    • Deployment Tracker
    • Funding Tracker
  • Deployments
  • Technology
  • Funding
  • Markets
  • Services
Humanoid Analytics
Home»Technology»Figure’s Helix 2.5 Brings Humanoid Robotics Closer to Its ChatGPT Moment
Technology

Figure’s Helix 2.5 Brings Humanoid Robotics Closer to Its ChatGPT Moment

Figure’s 30-home experiment suggests broad human-data pretraining may be starting to transfer across unfamiliar physical environments, but reliability and independent validation remain important gaps.
By Rinat MirzaitovSeptember 18, 20269 Mins Read
Image source: Figure AI.
Share
LinkedIn Twitter Copy Link Email

Humanoid robotics may be getting close to its ChatGPT moment, but not for the reason another polished robot video has gone viral.

The more important signal is underneath Figure’s new Helix 2.5 demonstration. Figure says a humanoid pretrained on a large dataset of human behavior was able to take three learned household behaviors into 30 homes it had never seen, using unfamiliar objects and without collecting new training data in those homes. In a controlled comparison disclosed by Figure, Index pretraining increased full-task zero-shot success from 9% to 56%.

That is still far from a dependable household robot. A 56% completion rate means roughly 44% of evaluated attempts did not successfully complete the entire task under Figure’s criteria. Independent reporting based on Figure’s disclosed evaluation counted 237 successful trials out of 420, including 94 of 140 bed-making trials, 87 of 140 towel-folding trials and 56 of 140 living-room tidying trials.

But reliability is only one part of what Figure is testing.

The bigger question is whether robotics is beginning to discover the same underlying development loop that transformed large language models: broad pretraining, followed by increasingly predictable improvements as data and compute scale.

If that mechanism transfers to physical intelligence, the commercial implications could be much larger than any individual household chore.

The Important Result Is Transfer, Not Bed Making

Figure’s accompanying video focuses on three visually understandable tasks: tidying a living room, making a bed and folding towels. The supplied Figure transcript describes robots entering homes they had never visited and manipulating objects the company says were absent from the task-specific training data. Figure says it observed successful behavior in every one of the 30 homes.
That language requires context. “Success in every home” does not mean every trial succeeded. Figure’s fuller evaluation reports 56% full-task success across the zero-shot evaluation.

More importantly, “zero-shot” does not mean Helix 2.5 was given completely unfamiliar chores and spontaneously worked out how to perform them.

Figure says the three behaviors were specified using fine-tuning data collected elsewhere. Zero-shot refers to the evaluation homes and manipulated objects. No data was collected in the 30 evaluation homes, Figure says, and a single fixed checkpoint for each task was used across all of them without adapting weights to individual homes.

That distinction prevents the result from being overstated. This is not yet the physical equivalent of typing an arbitrary new request into ChatGPT.

It is still commercially important.

Humanoid robots will struggle to scale if every factory, warehouse or home requires a new data-collection campaign and environment-specific model training. The cost of deployment could remain dominated by engineering rather than manufacturing.

A model that learns a behavior once and transfers it across many materially different environments attacks that bottleneck directly.

Humanoid Analytics previously highlighted this issue in Figure’s refrigerator experiments, where the important question was whether learning from broader physical experience could improve an apparently separate task:

Figure’s Fridge Story Shows Why Transfer Learning May Matter

Helix 2.5 is a stronger test of the same thesis because Figure now reports a controlled pretraining comparison and a multi-environment evaluation rather than a single anecdotal example.

This Is Where the ChatGPT Analogy Becomes Interesting

The “ChatGPT moment” analogy becomes useful only when it describes a change in the way capabilities are created.

Traditional automation is mostly built task by task. The system is engineered around a defined workflow, environment and set of edge cases.

Foundation models changed software AI because developers discovered that large-scale pretraining could create capabilities that transferred far beyond the individual examples in the training set. New applications could then be built on top of a broadly capable base model.

Figure is making the equivalent bet for robotics.

Its Index system collects human physical behavior at scale. Figure said in August that the platform was processing about 30 minutes of uploaded video every second. By the Helix 2.5 announcement, the company said that figure had reached roughly 35 minutes of new human experience per second.

The important Helix 2.5 experiment then attempted to isolate whether that pretraining actually mattered.

Figure says it trained two policies using identical task-specific data while holding architecture, optimization, hyperparameters and evaluation constant. One started from random weights. The other started from the Index-pretrained Helix 2.5 model.

The policy without Index pretraining achieved 9% full-task zero-shot success. The Index-pretrained version achieved 56%, according to Figure.

That does not independently prove Figure’s interpretation. The experiment was designed, executed and reported by Figure.

But it is a materially more useful piece of evidence than another capability demonstration because the company is attempting to identify what created the improvement.

Figure reported a second result that makes the language-model analogy more consequential. Across four training runs spanning an eightfold increase in Index pretraining data, the company says robot-action prediction loss improved predictably as data increased. Using the smaller runs, Figure says it forecast the largest run’s test loss to four decimal places, with forecasting error equal to 0.54% of the variation across the tested data range.

This is not proof that useful real-world robot performance will scale indefinitely with data. Figure held model size and downstream training fixed, and the reported scaling relationship concerns action-prediction loss rather than household task reliability.

Still, this may be the most important part of the announcement.

If developers can increasingly predict how robotic models improve when physical training data expands, humanoid development starts to look less like a collection of bespoke robotics projects and more like a scalable machine-learning program.

The Economics Could Change Before the Robot Is Perfect

For investors, that distinction matters because generalization affects the marginal cost of deployment.

A humanoid that requires extensive retraining every time it enters another customer facility may have an attractive hardware bill of materials and still produce poor economics. Engineering labor, data collection, teleoperation, integration and model tuning can consume the supposed flexibility advantage of a general-purpose machine.

Pretraining potentially changes that equation.

Figure says Helix 2.5 matched the success rate of a comparable Helix 02 behavior while using half as much task-specific adaptation data, then transferred the newer behavior across 30 unseen environments.

If that relationship survives broader testing, the important unit of scale is no longer merely the number of robots manufactured. It becomes the amount of useful physical experience that can improve future robot behavior across the fleet.

That also explains Figure’s extraordinary appetite for compute. Earlier this month, Figure and Nscale announced an initial $3.5 billion compute commitment involving access to as many as 100,000 Nvidia Vera Rubin GPUs, with initial deployment targeted for the second half of 2027. Figure explicitly connected that infrastructure to scaling Helix.

The commitment is evidence of Figure’s strategy and capital intensity, not evidence that the strategy will succeed. More compute does not increase a Humanoid Analytics Evidence Score and does not establish commercial deployment.

But the combination of Index, Helix and large-scale compute makes Figure’s thesis increasingly clear: collect human physical experience at massive scale, convert it into a pretrained physical model, then amortize that learning across robots and environments.

That is recognizably the foundation-model playbook.

The ChatGPT Moment Has Not Arrived Yet

There are at least three important gaps between Helix 2.5 and a genuine mass-market inflection.

The first is reliability. A 56% complete-task success rate is a research result, not a dependable service level. Homes are particularly unforgiving environments because robots eventually need to operate around people, pets, fragile objects and unpredictable interruptions.

The second is breadth. Helix 2.5 generalized three previously specified behaviors into new environments. A stronger foundation-model result would show one system accepting many new tasks with much less task-specific training, ideally through ordinary language or demonstration.

The third is independent operating proof. The 30-home evaluation was run by Figure. The homes were unfamiliar to the model, but they were evaluation environments rather than paying customer deployments. The available evidence therefore supports Internal Testing, not Operational Deployment under Humanoid Analytics methodology. Figure’s disclosed methodology is unusually useful for assessing the result, but independent reproduction, intervention data, repeated operation and customer evidence would materially strengthen the conclusion.

This is why the ChatGPT analogy should not be attached to the sight of a humanoid folding a towel.

The real moment will come when a broadly pretrained robot can enter unfamiliar human environments, perform a growing range of useful work without site-specific retraining, improve predictably as its training base expands, and do so reliably enough that deploying another robot becomes primarily an operational decision rather than another robotics research project.

Helix 2.5 does not establish that endpoint.

It does, however, provide evidence that one of the mechanisms required to reach it may be starting to work.

The next decisive result is not another perfect demonstration. It is a higher success rate across more tasks, repeated in environments Figure does not control, with disclosed intervention and operating data.

If that evidence begins to appear while the underlying scaling curve continues, humanoid robotics will have moved much closer to something the software AI industry has already experienced: the point where a general model stops looking like an interesting research system and starts looking like a platform.

That is the moment worth watching.

Sources:

  1. Figure AI, “Helix 2.5: Zero-Shot 30-Home Generalization”
    Source type: Tier 3, detailed first-party disclosure; company-controlled
    Figure AI source
  2. Figure, “Helix 2.5 30-Home Generalization”
    Source type: Tier 4, company-controlled promotional and demonstration video
    https://www.youtube.com/watch?v=lJpM_2a1zrE
  3. Humanoids Daily, “Figure’s Helix 2.5 Takes on Chores in 30 Unseen Homes”
    Source type: Tier 2, independent reporting based primarily on Figure’s disclosed experiment and footage; not independent reproduction
    Humanoids Daily source
  4. Figure AI, “Introducing Index: Building The World’s Largest and Most Diverse Physical Dataset”
    Source type: Tier 3, detailed first-party disclosure; company-controlled
    Figure AI Index source
  5. Figure AI, “Figure and Nscale Sign Strategic Partnership For Up to 100,000 GPUs on the NVIDIA Vera Rubin Platform”
    Source type: Tier 3, detailed first-party disclosure; company-controlled
    Figure and Nscale source
Featured Market Signals Partially Confirmed Claim Selected Analysis
Share. LinkedIn Twitter Copy Link Email

Related Analysis

Agility’s Digit 5 Adds Scale Features, but Operating Proof Is Still Ahead

September 16, 2026

OpenAI and Anthropic Enter the Humanoid Robotics Race From Opposite Ends

September 3, 2026

Standardizing Physical AI: Why Common Hardware Interfaces Could Accelerate Generalization

August 31, 2026

Why Humanoid Robotics Is Hard: Telemetry, Physical AI, and the Data Loop

August 11, 2026
Selected Analysis

Agility’s Digit 5 Adds Scale Features, but Operating Proof Is Still Ahead

September 16, 2026

China Raises the Reported IPO Bar for Humanoid Robotics

September 12, 2026

China Controls 90%+ of Some Critical Components Inside a Humanoid Robot

September 10, 2026

Humanoid Analytics tracks the commercial progress of humanoid robotics through evidence-based analysis, company profiles, deployment trackers, and market intelligence.

Connect with us:

X (Twitter) Mastodon LinkedIn RSS
Highlights

Figure’s Helix 2.5 Brings Humanoid Robotics Closer to Its ChatGPT Moment

September 18, 2026

Agility’s Digit 5 Adds Scale Features, but Operating Proof Is Still Ahead

September 16, 2026

China Raises the Reported IPO Bar for Humanoid Robotics

September 12, 2026
Stay Informed

Subscribe to Updates

Get evidence-based updates on humanoid robotics companies, deployments, funding, partnerships, and market signals.

  • Home
  • About Us
  • Our Services
  • Evidence Standards
  • Contact Us
  • Privacy Policy
  • Terms of Use
© 2026 Humanoid Analytics. All rights reserved.

Type above and press Enter to search. Press Esc to cancel.