The Dirty Data Problem That Is Stalling the Next Wave of AI Robots
A new survey of more than 700 AI engineers finds that bad data, not weak algorithms, is the main reason physical AI systems fail before they reach the real world.

Key points
- A 2026 survey of more than 700 AI professionals, reported by IEEE Spectrum AI, finds data quality problems cause the majority of failures in physical AI systems.
- 78% of teams already see measurable results from visual and physical AI, yet 74% believe the field receives far less investment than it deserves.
- Teams that successfully ship physical AI products spend nearly three times as long preparing and curating data as teams whose projects stall.
- 92% of practitioners agree on where the field is heading next, signalling strong consensus about what physical AI will need to do.
For the past decade, the biggest AI breakthroughs, large language models, image generators, voice assistants, were built on text scraped from the internet. Now the frontier has shifted to the physical world.
Systems that drive cars, guide robotic arms, or steer delivery drones don't read sentences. They read video feeds, LiDAR scans (laser-based 3D maps of the surroundings), and sensor data arriving thousands of times a second. Building AI that acts in physical space is genuinely harder than teaching it to write an email.
So what is actually going wrong?
Data problems cause most failures, not weak AI models. A 2026 survey of more than 700 professionals, conducted with support from Voxel51 and covered by IEEE Spectrum AI, found that teams are collecting vast amounts of video and sensor data but struggling to turn that raw material into something useful.
The clearest bottleneck is annotation: the slow, expensive process of labelling what appears in each frame of footage so the AI can learn from it. Teams often label everything they collect, then discard most of it before the system ships. That wasted effort eats budgets and slows timelines. Our 5 August story on industrial perception, "The Robot That Can Move Is Useless If It Cannot See", found a mining robotics company making exactly this point: the data pipeline, not the motion controller, was where projects broke down.
What separates teams that ship from teams that stall?
Time spent on data work, not better algorithms. Successful teams invest nearly three times as long curating and selecting their training data compared with teams whose projects never reach production.
That's a striking result. The instinct in AI development is often to chase a bigger model or a newer architecture. This survey says the smarter bet is to be ruthless about what you feed the model in the first place.
| Finding | Figure |
|---|---|
| Teams seeing measurable value from physical AI | 78% |
| Teams saying the field is underinvested | 74% |
| Data work advantage for successful teams | ~3x more time |
| Practitioners agreeing on where the field heads next | 92% |
What does this mean for ordinary people?
Physical AI is already around us: warehouse robots that sort packages, hospital robots that carry supplies, self-driving software in new cars. These systems will arrive, but unevenly, because building them well is unglamorous, labour-intensive work.
The 74% who consider the field underinvested may well be right, and the implication is practical: if resources flow into data infrastructure rather than headline-grabbing model launches, the physical AI products consumers actually use could become meaningfully more reliable.
The message from the people building these systems is blunt. The algorithm's rarely the problem. The data is.



