Figure Releases Index Dataset: 16 Million Crowdsourced Videos Reshape Robot Training

Figure.AI launches Index, a 16-million-video crowdsourced dataset to train general-purpose humanoid robots.
Figure.AI has released Index, billed as the largest and most diverse robot training dataset ever, containing 16 million videos collected through an open crowdsourcing model where anyone can record daily tasks and earn money. The dataset aims to solve the critical data scarcity problem in embodied AI, providing the diverse real-world data needed to train robots that can generalize across tasks and environments — a key step toward truly general-purpose humanoid robots.
Figure Launches Index: The Largest Robot Dataset in History
The humanoid robotics field has reached another major milestone. Figure.AI has officially released a robot training dataset called Index — billed as the largest and most diverse robot dataset ever created, containing a staggering 16 million videos. This figure far exceeds most previously published robot learning datasets, signaling that humanoid robot training has officially entered the "big data" era.

For the "data hunger" problem that has long plagued robot learning, Index attempts to offer a scalable answer. Unlike autonomous driving and large language models, which have long enjoyed the benefits of massive datasets, the robotic manipulation domain has consistently lacked unified, large-scale, high-quality real-world data. Previously, the most influential robot datasets — such as Google's RT series data and the Open X-Embodiment project — typically ranged from hundreds of thousands to a few million trajectories, and were often limited to specific robot platforms and a narrow set of task types. By comparison, autonomous driving datasets like Waymo Open Dataset and nuScenes have accumulated millions of miles of driving data, while LLM training corpora are measured in trillions of tokens. The cost of collecting robotic manipulation data is far higher than for text and images — every data point requires physically executing an operation in the real world. This is the fundamental reason the field has long faced data scarcity. Figure's move is designed precisely to close this gap.
Why Robot Training Needs Such a Massive Dataset
Data Is the Core Bottleneck for Embodied Intelligence
For humanoid robots to truly enter homes and factories, the core challenge isn't the hardware itself — it's the "brain," meaning manipulation policies that can generalize across diverse tasks and environments. Training such policies requires covering sufficiently diverse scenarios: different objects, lighting conditions, household layouts, and manipulation methods.
A key concept to understand here is Embodied AI — a technology paradigm that embeds AI systems within physical entities (such as robots), enabling them to perceive, understand, and interact with the real physical world. Unlike AI operating purely in the digital realm, the core challenge for embodied intelligence is the "sim-to-real gap" — manipulation policies trained in simulated environments often fail to transfer directly to the real world. The physical properties of the real world are extraordinarily complex: subtle variations in friction, nonlinear deformation of soft objects, random fluctuations in lighting conditions — all of which current physics simulators struggle to perfectly reproduce. Therefore, large-scale real-world data has become the most direct path to bridging this gap.
Data collected in a single lab tends to be highly homogeneous, and models trained on it experience sharp performance drops once they leave controlled environments. The diversity that Index emphasizes directly addresses this pain point. If 16 million videos can truly cover a rich variety of everyday task scenarios, they will provide an unprecedented data foundation for training manipulation models with generalization capabilities.
A Critical Step from "Specialized Robots" to "General-Purpose Robots"
The industry widely believes that general-purpose humanoid robots need a "scaling effect" similar to that of large language models — once data volume and diversity reach a critical threshold, model capabilities undergo a qualitative transformation. This view stems from the Scaling Law systematically articulated by OpenAI in 2020: the performance of large language models follows a power-law relationship with model parameter count, data volume, and compute — when these factors scale simultaneously, model capabilities improve continuously and predictably. Since 2023, work from Google DeepMind such as RT-2 and Octo has begun to validate that similar scaling effects exist in the robotics domain — when training data covers more robot morphologies, task types, and environments, zero-shot generalization capabilities improve significantly. However, whether the robotics field has a similar "emergent ability" threshold like language models remains an open scientific question.
Figure is betting on exactly this logic: first build an unprecedentedly large data foundation, then leverage LLM-style training paradigms to approach general-purpose manipulation capabilities.
Crowdsourcing Model: Anyone Can Record, Anyone Can Get Paid
A New Paradigm for Robot Data Collection
The most noteworthy innovation in Index is its open crowdsourcing collection mechanism. According to the release information, starting now, anyone can record their daily tasks and get paid for it.
This design is quite disruptive. Traditional robot data collection relies on professional teams, expensive equipment, and controlled environments — it's costly, inefficient, and limited in scope. The crowdsourcing model extends the collection network into millions of households, allowing real-world diversity to flow naturally into the dataset — different family kitchens, different people's manipulation habits, different regional lifestyles can all become valuable material for robot training.
It's worth noting that this technical approach isn't entirely without precedent. In the robotics field, there have been similar crowdsourcing attempts: Google's ALOHA project uses low-cost teleoperation devices to collect bimanual manipulation data, and the DROID dataset collected 76,000 manipulation trajectories across 13 research institutions. But Figure's Index far surpasses these predecessor projects in both scale and openness. Technically, the key challenge for crowdsourced robot data lies in standardizing the collection protocol — different users use different cameras, different shooting angles, and different resolutions. Data consistency and usability require carefully designed collection tools and post-processing pipelines.
The Dual Significance of Incentive Mechanisms
"Record and earn" isn't just a data acquisition method — it's a sustainable ecosystem design. Through economic incentives, Figure can continuously expand its data scale at relatively low marginal cost, while enabling ordinary users to become both participants in and beneficiaries of AI progress.
This echoes the approach of early crowdsourcing annotation platforms like Amazon Mechanical Turk. Launched in 2005, Amazon Mechanical Turk pioneered large-scale human annotation — the famous ImageNet dataset was annotated through this platform, with 14 million images manually labeled, directly sparking the deep learning revolution of 2012. But Index's crowdsourcing target has been upgraded from static image annotation to dynamic embodied task demonstrations. Each data point contains spatiotemporal continuous information about human manipulation (hand motion trajectories, force application patterns, task decomposition logic, etc.), delivering far higher value density than traditional annotation data.
Of course, this model also introduces a series of challenges around data quality control, privacy protection, and content filtering. How to strike the right balance between scale and quality will be key to Index's success.
Index's Potential Impact on the Robotics Industry and Unanswered Questions
Profound Implications for the Embodied Intelligence Industry
If Index's model proves viable, it could reshape the entire robot data supply chain landscape. Data would no longer be a scarce resource monopolized by a few tech giants, but rather a nascent public infrastructure formed through distributed crowdsourcing. This has profound implications for accelerating the entire embodied intelligence field.
At the same time, this move reinforces Figure's differentiated positioning in the humanoid robot race — it's not just about building hardware, but also about commanding the soft-power moat of "data + models." Figure.AI was founded in 2022 by Brett Adcock and is one of the most closely watched startups in the humanoid robotics space. In early 2024, Figure completed a $675 million funding round with participation from Microsoft, NVIDIA, OpenAI, and Jeff Bezos, reaching a valuation of $2.6 billion. Its flagship product, Figure 02, has established partnerships with manufacturers such as BMW. In terms of the competitive landscape, Figure faces rivals including Tesla Optimus, Boston Dynamics Atlas, 1X Technologies' NEO, and Agility Robotics' Digit. In this race, data and model capabilities are increasingly becoming more critical differentiators than hardware — whoever commands the richest real-world manipulation data may be the first to train a truly general-purpose robot brain. This is the deeper strategic logic behind Figure's heavy investment in the Index dataset.
Several Key Questions That Warrant Sober Assessment
Interestingly, the information available so far comes primarily from community discussions, and Figure has yet to disclose further official technical details. Several key questions about Index remain to be answered:
- Data Quality: What are the specific sources, annotation methods, and quality distribution of the 16 million videos? Crowdsourced data inherently contains noise and inconsistencies — has Figure deployed automated quality screening and cleaning pipelines?
- Openness: Is this data fully open, or reserved solely for Figure's internal training? If open, under what license will it be released?
- Privacy Compliance: Crowdsourced recordings of daily tasks involve household privacy — how will user data security be ensured? Especially as data protection regulations like GDPR become increasingly strict, video collection involving home environments faces higher compliance requirements.
- Actual Effectiveness: Can the data scale truly translate into significant leaps in robot manipulation performance? Will increasing data volume encounter diminishing marginal returns?
Until these details are clarified, assessments of Index should maintain rational optimism.
Conclusion
Regardless of its ultimate results, Figure.AI's Index dataset represents an important direction for robot learning: using scalable, crowdsourced real-world data to tackle the generalization challenge of embodied intelligence. When "everyone can contribute data to robot training and get paid for it" becomes reality, we may be standing at the turning point where humanoid robots transition from the laboratory to everyday life. Going forward, it will be worth closely following the additional technical details and real-world deployment results that Figure releases.
Related articles

MTNode 1.2.4 Update Explained: App Slimming, Bug Fixes, and Differential Algorithm for Transparent Channel Generation
MTNode 1.2.4 brings three key improvements: canvas deletion bug fix with backup recovery, app slimming for faster installs, and a differential algorithm for generating transparent channels in AI images.

Speechmark: A Fully Offline Mac Meeting Transcription Tool That Keeps All Data on Your Device
Speechmark is a privacy-first macOS meeting transcription tool. Recording, transcription, and summarization all happen locally with no cloud uploads required.

hob: A Professional AI Workbench for Managing Multi-Agent Collaboration
hob is a professional workbench for the AI Agent stack, unifying multi-model orchestration, workflow automation, review, and recovery in one interface for managing multi-Agent collaboration.