Figure AI's New Data Platform Pays Millions to Crowdsource Robot Training Videos
Figure AI has publicly launched Index, a consumer-facing app that crowdsources real-world human video to train its Helix AI architecture for robots. The platform, which emerged from stealth after nearly a year of quiet operation under the codename Project Go-Big, has already amassed 16 million video uploads from more than 44,000 weekly active users across 108 countries, processing 30 minutes of footage every second.
Why Does Figure Need Crowdsourced Video Data?
Robotics has long faced a critical bottleneck that sets it apart from large language models. While AI systems like ChatGPT can train on the vast corpus of publicly available text on the internet, robots need something fundamentally different: real-world physical interaction data showing how humans manipulate objects, navigate spaces, and complete tasks. Traditional methods for collecting this data, such as direct robot teleoperation or kinesthetic teaching, remain slow, labor-intensive, and difficult to scale across varied environments.
Figure's answer is to turn everyday people into data collectors. Index operates as a two-sided platform where individual "Creators" can record themselves performing household or workplace tasks for cash bounties, or users can book gig workers through the app to complete chores on-site while capturing first-person footage. This approach bypasses the traditional data bottlenecks by tapping into a global workforce willing to generate embodied interaction data at scale.
What Scale Has Index Already Achieved?
During its four months operating in stealth, Index demonstrated substantial traction before its public launch. The platform recorded 264,000 app downloads across 108 countries, with 44,000 or more weekly active users contributing data. The sheer volume is striking: 16 million video uploads translating to 30 minutes of physical video ingested every second, or roughly 4.9 years of human work uploaded per day.
The diversity of the dataset is equally important. For every 1,000 hours of video logged, Figure reports that the dataset captures 373 unique tasks, 1,146 distinct manipulated objects, and 116 unique environments. Captured workflows span domestic chores such as folding laundry, making beds, and cleaning, to commercial operations in restaurants, retail stockrooms, and logistics facilities.
How Does Figure Ensure Data Quality at This Scale?
Ingesting continuous consumer-generated video at high volume introduces major quality-control challenges. To convert raw smartphone video into structured training data for the Helix model, Figure built an automated, five-stage ingestion pipeline designed to filter noise and extract meaningful patterns.
- Filtering: Automated semantic and technical vision filters immediately screen uploads for resolution, frame rate, lighting conditions, and task relevance.
- Fraud Review: Dedicated human auditing teams monitor user accounts to detect deliberate attempts to spoof tasks or game payout incentives.
- Deduplication: High-dimensional vector embeddings are generated for each video segment to identify and discard redundant footage that falls above strict similarity thresholds.
- Rebalancing: The dataset is dynamically balanced using task quotas and embedding clusters to prevent over-representation of simple tasks while prioritizing rare edge cases.
- Hierarchical Annotation: Accepted episodes receive structured hierarchical text captions, aligning high-level task descriptions with low-level physical manipulations for multimodal model training.
This multi-stage approach ensures that Figure is not simply accumulating video, but systematically converting it into high-quality training material that can teach robots to generalize across diverse physical scenarios.
What Financial Commitment Is Figure Making?
Figure has already distributed $15 million in payouts to contributors and committed over $1 billion toward data acquisition and compute over the next 12 months. This represents a significant bet that human video alone can provide the foundation for general-purpose robot reasoning. Whether human video can capture the tactile feedback and force dynamics needed to master intricate physical manipulation remains an ongoing research question, but Figure is putting substantial capital behind the bet.
"Index serves as a dedicated pipeline to feed the company's Helix AI architecture with the rich, diverse physical interactions required for zero-shot generalization," explained Brett Adcock, founder and CEO of Figure AI.
Brett Adcock, Founder and CEO at Figure AI
How Does This Fit Into Figure's Broader Commercial Strategy?
The launch of Index marks a critical shift in Figure's broader commercial strategy. The company has already begun manufacturing hardware at scale, recently celebrating its 1,000th Figure 03 build at its BotQ facility and deploying units into real-world commercial pilots with BMW at Plant Spartanburg and Catalyst Brands in Reno.
However, as Adcock has argued, hardware scalability is no longer the primary hurdle. The true bottleneck is onboard intelligence and general-purpose reasoning. By scaling Index into a multi-million-dollar data collection engine, Figure is attempting to build the equivalent of ImageNet, the foundational dataset that accelerated computer vision, but for embodied AI. If successful, Index could become the training ground that transforms Figure's robots from specialized machines into genuinely adaptable systems capable of learning new tasks from diverse real-world scenarios.