Humanoid Robots Are Getting Smarter. They Are Not Ready for Your Home Yet.
Nvidia's Cosmos 3 gives robot builders a powerful new tool. But researchers warn that the gap between a robot handing out popcorn and one folding your laundry is still vast.

Key points
- Nvidia's Cosmos 3, a foundation model (a large AI system trained to understand the physical world) designed for robotics, is now available for developers building robot software.
- Google DeepMind's Gemini Robotics, a vision-language-action model that lets a robot see and act on what it understands, represents the current frontier: it can pack a lunchbox, but only if it has seen similar tasks before.
- Morgan Stanley estimates the humanoid robot market could reach $5 trillion by 2050, though leading AI researchers say the machines are nowhere near useful enough to justify that timeline.
- The hardware is advancing faster than the intelligence running it, a pattern that has held across all 35 robotics stories AI2Day has published in the past 30 days.
The white-and-black robot you may have seen handing out water bottles at a Tesla event, or struggling to iron a shirt, is Tesla's Optimus. Elon Musk says it could automate almost all human labour for around $20,000 a unit. Jensen Huang, CEO of chip-maker Nvidia, said in January that humanoid robots would match human ability this year. Both men are talking about a technology that, right now, can't reliably iron a shirt.
That gap between the promise and the product is exactly what Nvidia is trying to close with Cosmos 3, its new open physical AI foundation model. Think of it as a shared brain that robot builders can start from, rather than building from scratch. It can act as a vision-language model (software that understands both images and words) to reason about objects and intent in the real world. Developers can also use it to simulate entire environments before a robot ever touches a real object, running hundreds of possible outcomes in a virtual space to find the best behaviour.
The practical upside is speed. Teams building a warehouse robot no longer need to train a model from zero. They take Cosmos 3, feed it data specific to their robot's cameras and body, and the system adapts. We've tracked Nvidia's simulation work before: our September story on Nvidia's Warp showed how parallel GPU environments are already cutting the time it takes to teach robots physical skills.
How far along is the actual science?
Not as far as the headlines suggest. The best demonstration of where robot AI genuinely stands comes from Google DeepMind's Gemini Robotics, which controls a pair of mechanical arms called ALOHA 2. Reported in depth by MIT Technology Review, the system can pack a lunchbox, pick up snow peas with tongs, and do basic origami. Three years ago, none of that was possible. Today it's impressive. It's also brittle: the robot only succeeds at tasks it has already seen demonstrated, usually by a human operating it by remote control to generate training data.
This is the core problem. The AI methods that made chatbots fluent in language are being applied to physical movement, but the physical world is unforgiving in ways that language isn't. Miss a word in a sentence and a reader fills the gap. Miss a target by two centimetres and the cup stays on the table.
Yann LeCun, one of the scientists who laid the groundwork for modern AI, said bluntly in January: none of the companies building humanoid robots "has any idea how to make those robots smart enough to be useful."
What does this mean for ordinary people?
Not much, yet. Factory floors will see robot assistants first, and even there the rollout will be slow and task-specific. A robot dog patrolling a warehouse at night (a use case we covered on 7 October) is a very different machine from one that can load a dishwasher on demand.
For anyone worried about job displacement: the near-term threat is narrower than the headlines imply. Robots are getting genuinely better at fixed, repeatable tasks in controlled spaces. Unpredictable environments, the kind every human worker navigates every day, remain out of reach.
Tools like Cosmos 3 are real progress. They lower the cost of entry for robotics teams and will likely speed up the pace of improvement. A faster path to a difficult destination is still a long road, though.
Common questions
Will robots replace workers soon?
Not broadly. Current systems handle specific, repetitive tasks well but struggle badly with anything that requires adapting to a changing environment. Widespread job displacement in varied roles is at minimum a decade away, and most researchers think that timeline is optimistic.
What is a foundation model and why does Cosmos 3 matter?
A foundation model is a large AI system trained on huge amounts of data that other developers can build on top of, instead of starting from scratch. Cosmos 3 is built specifically for physical AI, meaning robots and autonomous vehicles, which makes it a practical shortcut for teams that would otherwise spend years gathering and processing training data.
Do I need to worry about what Nvidia is building?
Not directly. Cosmos 3 is a tool for companies building robots, not a product you buy. Its effects will show up gradually, as the robots those companies ship get incrementally better at their jobs.



