Physical AI has a data problem. To operate in the real world, models need to learn from human experience.
The scale of what that demands is increasingly clear:
↳ $7T potential humanoid market by 2050
↳ $11.9B → $49.8B AI training dataset market by 2031
↳ ~100K hours of
Data in the Wild, ep. 001.
The data needed to train robotics is rare, unscripted, and full of long-tail behaviors. Here’s what that looks like in practice.
Task context:
•- Brush the pet, wipe its paws & body, collect loose fur
•- 11 minutes
•- A home in the Philippines
Numo is now mobile-first.
Contribute from anywhere with new video tasks, audio tasks in more languages, and faster withdrawals.
Live today on iOS and Android ↓
Conversational voice data exposes a subtle failure mode in many ASR pipelines. The metric says the data is bad, but human reviewers say the audio is clear.
Often the issue isn’t transcription quality itself, but that the evaluation stack wasn’t built for turn-taking, silence,
A fascinating look at the challenge ahead for physical AI.
Exactly what we're building: the pipeline to collect high-quality real-world data at the scale physical AI demands.
World Labs CEO Dr. Fei-Fei Li & SceniX Co-Founder Yunzhu Li on the data bottleneck in robotics and how world models help:
"The lack of data in training, the lack of data in evaluation, this is very, very different from language models, where data is abundant on the internet."