Humanoid robotics is developing a training-data market, but it is not yet a commodity market.
Robot developers are building dedicated collection systems, specialist companies are offering physical AI data services, open datasets are distributing robot trajectories at scale, and early marketplaces are testing whether robotics datasets can be verified, licensed and sold.
The demand is real. The less certain question is where durable commercial value will sit.
That question matters for another reason. Data is not simply an input purchased by robot developers. For learned robotic systems, it is part of the path toward greater autonomy.
A humanoid operating outside a tightly controlled workflow has to cope with different objects, positions, environments, instructions and failure cases. Developers are trying to build models that can transfer what they have learned rather than requiring a new set of manually engineered behaviors for every task.
That makes the quality and diversity of physical experience strategically important.
For investors and strategy teams, the commercial question is therefore broader than how many hours of robot data the industry may require. It is whether those hours create a large independent supply market, whether the most valuable data remains proprietary, and which companies can demonstrate that their data actually produces better physical AI.
Why Data Matters for Humanoid Autonomy
Traditional industrial automation works well when machines repeat carefully engineered actions in structured environments. General-purpose humanoids face a different problem.
A robot expected to operate across homes, warehouses, factories or service environments cannot realistically have every object position, grasp, obstruction and failure mode programmed in advance. Learned policies instead use data to develop patterns that can be applied across variations in tasks and environments.
This is where data becomes central to physical AI.
Google DeepMind’s Open X-Embodiment project combined robotics data from more than 20 institutions and multiple robot types. The underlying research direction is important because it tests whether experience collected across different embodiments and tasks can contribute to more general robot policies rather than isolated task-specific systems.
Physical Intelligence is pursuing a related approach. Its π0 work combines internet-scale pretraining, open robotics datasets and proprietary data collected across multiple robots and dexterous tasks. The objective is a generalist policy capable of using broad prior experience and then adapting to downstream tasks.
These are research programs, not proof that general-purpose humanoid autonomy has been solved.
Data is also not sufficient by itself. Model architecture, sensing, control, compute, hardware reliability and safety all remain important. A large dataset cannot compensate automatically for poor hardware, unsuitable sensors or a model that cannot use the information effectively.
But the data problem is fundamental because autonomy requires robots to deal with variation.
The commercially meaningful evidence would therefore not be a headline number of videos or trajectories. It would be measurable improvement in areas such as task success across unfamiliar conditions, recovery from failures, adaptation to new objects and environments, reduced human intervention and the amount of additional task-specific training required for a new workflow.
That distinction matters when evaluating the emerging data race. More data is a resource. Better generalization is the result developers ultimately need to prove.
Robot Developers Are Building Their Own Data Loops
Figure AI provides one of the clearest examples of a company treating data collection as strategic infrastructure.
In August 2026, the company introduced Index, a physical-data collection system. Figure said the platform had accumulated more than 16 million uploaded videos and paid creators more than $15 million. It also said previous attempts to buy external data did not provide the throughput, diversity or quality required for its Helix models.
Those are company-reported figures rather than independently audited operating metrics.
Figure also explicitly presents generalization as a data problem. That is a company position rather than independent evidence that Index has produced superior autonomy, but it helps explain why Figure is investing in a proprietary collection engine rather than relying only on external datasets.
1X is taking a similar approach. Its World Model Lab describes a mixture of internet-scale media, egocentric human video, simulation, remotely operated robot data and information generated by its own NEO humanoids. The company presents control of this learning loop as a competitive advantage.
Apptronik is also building physical collection infrastructure. Its expanded Robot Park in Austin is described as a nearly 90,000-square-foot facility where Apollo humanoids generate training data, including work linked to its research relationship with Google DeepMind.
Physical Intelligence offers a more hybrid model, combining open robotics datasets and internet-scale pretraining with proprietary robot interaction data.
Together, these examples suggest that major developers may use several data sources while keeping strategically important collection and post-training data inside the company.
That limits a simple thesis that rising robot-data demand automatically creates an equally large third-party market.
If proprietary operating experience improves a developer’s models, data can potentially become part of a feedback loop. Robots generate experience, that experience contributes to model development, and improved models may then produce additional useful operating data.
Whether any company has already created a durable version of that loop at commercial humanoid scale remains unproven publicly. But the incentive to own it is increasingly visible.
External Supply Is Expanding From Services to Marketplaces
Independent supply is developing through several models.
Scale AI markets a Physical AI Data Engine built around robotics data factories, distributed collectors and real operating environments. The company says its network can collect more than 1,000 hours of demonstration data per day across formats including teleoperated demonstrations and egocentric human data.
That is a first-party capacity claim, not evidence that every collected hour produces useful model improvement.
Instawork approaches the problem through access to workers and workplaces. Its robotics offering includes egocentric video, multimodal human demonstrations and teleoperated robot-learning data.
The potential advantage in this model is not simply annotation capacity. Access to people performing relevant tasks, real workplaces and diverse physical environments may itself become scarce infrastructure.
Open datasets create another competitive force.
Google DeepMind’s Open X-Embodiment project brought together robotics data from more than 20 institutions. Hugging Face’s LeRobot ecosystem provides standardized tools for packaging and distributing multimodal robot-learning datasets.
Open data does not eliminate commercial demand, but it raises the bar for paid providers. A supplier increasingly has to explain why its data is scarcer, better licensed, higher quality, more task-specific or more useful than what developers can obtain without purchasing a proprietary dataset.
Synthetic data adds another variable. NVIDIA and other physical AI platforms are building simulation and synthetic-data workflows that can expand training sets without physically collecting every variation or failure case.
Synthetic generation could be particularly useful where physical collection is expensive, rare, dangerous or difficult to reproduce. But simulated volume is not automatically equivalent to useful real-world experience. The important question is how effectively models trained with synthetic information transfer into physical operation.
The emerging market therefore includes proprietary collection, managed collection services, open datasets, simulation and synthetic generation.
It also now includes marketplaces.
Kinetic Blocks is notable because it addresses discovery, verification, licensing and price formation rather than collecting all the underlying data itself.
The beta platform describes itself as an open marketplace for robotics training data. Public listings show characteristics such as hours, clips, modality, geography, verification status and quality grade. Kinetic Blocks says sellers retain ownership and can determine price, licensing and exclusivity terms.
The company also distinguishes between verifying that underlying files correspond to a listing and grading their technical and content quality.
That distinction matters. A functioning market requires buyers to trust both what they are purchasing and whether the dataset is useful for their models.
China’s RoboMIND provides another view of data becoming physical infrastructure.
A September 3, 2026 Global Times report said the Beijing Innovation Center of Humanoid Robotics reported that its open-source RoboMIND dataset had exceeded 20 million cumulative downloads, doubling within one month. The center said RoboMIND contained more than 300,000 dual-arm manipulation trajectories covering more than 700 real-world tasks.
The center also said its nearly 6,000-square-meter training facility operates more than 150 robots across 40 configurations and had delivered nearly 30,000 hours of data to external partners.
These figures indicate substantial collection and distribution if taken as reported, but they require careful interpretation.
Downloads demonstrate distribution, not improvement in physical AI. Hours delivered measure supply, not customer economics. Neither figure establishes that the data reduced interventions, improved task success or enabled a robot to generalize into an unfamiliar operating environment.
The broader signal is stronger than any individual number.
Physical AI data collection is increasingly being organized as infrastructure involving facilities, robots, operators, human demonstrations, environments, simulation and standardized data workflows.
What Would Prove a Durable Data Business
The market is likely to reward more than volume.
A commercially valuable dataset may need clear rights, difficult-to-reproduce tasks, diverse environments, useful variation, synchronized sensor streams, strong metadata and enough quality control that customers do not spend more cleaning and filtering the data than acquiring it.
For autonomy, another test becomes important: does the data contain experience that helps a model deal with situations beyond the narrow examples on which it was trained?
That could include variation in objects, layouts, human behavior, task sequence and failure conditions. Data capturing only repeated examples of an already solved behavior may be less strategically valuable than smaller datasets containing difficult, diverse or previously unseen situations.
The strongest commercial evidence for external suppliers would therefore combine two forms of proof.
The first is repeat customer purchasing. That would show that buyers consider the data valuable enough to acquire more than once.
The second is buyer-confirmed performance improvement. That could include evidence that externally supplied data improves task success, generalization, recovery, intervention rates or the amount of additional training required to deploy a model into a new workflow.
For large open datasets such as RoboMIND, downloads are useful evidence of interest and distribution, but downstream adoption and reproducible performance improvements would be more informative.
Vertical integration remains the main alternative outcome.
Figure and 1X both present proprietary data loops as strategic assets. If that pattern continues, independent providers may find their strongest markets in rare tasks, specialized environments, regional access, short-term collection surges and smaller robotics developers that cannot economically justify building their own data factories.
External suppliers could also become more valuable where they provide something developers cannot easily reproduce themselves: legal rights, geographic diversity, specialized workers, unusual environments, scarce failure cases or high-quality data collected on demand.
The humanoid training-data market is therefore real enough to matter, but too early to value simply by hours collected.
Its strategic importance may ultimately be larger than the dataset market itself. If more capable physical AI depends partly on learning from sufficiently broad and relevant physical experience, then the systems used to collect, verify, curate and improve that experience become part of the infrastructure behind humanoid autonomy.
But data volume is not autonomy.
The breakthrough evidence would be robots that can transfer learning across tasks, cope with unfamiliar conditions, recover when actions fail and operate with progressively less human intervention. For data companies, the equivalent commercial proof would be evidence that customers repeatedly pay for datasets that help produce those improvements.
The companies most likely to create durable value will therefore not necessarily be those with the largest headline dataset. They will be those that can prove their data is difficult to reproduce, legally usable, operationally relevant and measurably useful to the physical AI systems that consume it.
Sources:
- Google DeepMind, “Scaling up learning across many different robot types”
Source type: Tier 3, detailed first-party research disclosure, company-controlled
https://deepmind.google/blog/scaling-up-learning-across-many-different-robot-types/ - Physical Intelligence, “π0: Our First Generalist Policy”
Source type: Tier 3, detailed first-party technical disclosure, company-controlled
https://www.pi.website/blog/pi0 - Figure AI, “Introducing Index: Building The World’s Largest and Most Diverse Physical Dataset”
Source type: Tier 3, detailed first-party disclosure, company-controlled
https://www.figure.ai/news/introducing-index - 1X, “1X Launches World Model Lab to Scale Humanoid Intelligence”
Source type: Tier 3, detailed first-party disclosure, company-controlled
https://www.1x.tech/discover/1x-world-model-lab - Apptronik, “Welcome to Robot Park: Where Apptronik’s Apollo Goes to Work Training the Next Generation of Humanoid Robot Intelligence”
Source type: Tier 3, detailed first-party disclosure, company-controlled
https://apptronik.com/news-collection/welcome-to-robot-park-where-apptroniks-apollo-goes-to-work - Scale AI, “Physical AI | Scale AI”
Source type: Tier 3, detailed first-party product disclosure, company-controlled
https://scale.com/physical-ai - Instawork, “Robotics | Instawork”
Source type: Tier 3, detailed first-party product disclosure, company-controlled
https://www.instawork.com/robotics - Hugging Face, “LeRobotDataset v3.0”
Source type: Tier 3, official technical documentation, company-controlled
https://huggingface.co/docs/lerobot/lerobot-dataset-v3 - NVIDIA, “Synthetic Data for AI & 3D Simulation Workflows”
Source type: Tier 3, detailed first-party technical and product disclosure, company-controlled
https://www.nvidia.com/en-us/use-cases/synthetic-data-physical-ai/ - Kinetic Blocks, “Kinetic Blocks | The Open Marketplace for Robotics Training Data”
Source type: Tier 3, detailed first-party disclosure, company-controlled
https://kineticblocks.com/ - Global Times, “Beijing-based humanoid robotics open-source dataset hits 20m global downloads, doubling in one month”
Source type: Tier 2, independent reporting containing attributable first-party claims
https://www.globaltimes.cn/page/202609/1369689.shtml
