Robots Can Move. They Can See. So Why Can't They Hold a Conversation?

Humanoid robots are impressive to watch, until you try talking to one. A specialist in audio AI explains why poor hearing and bad voice interaction could be the thing that stops robots from ever being widely accepted.

AI2Day Newsdesk4 min read
A modern urban street at dusk photographed from street level, showing a sleek electric vehicle with a sensor array on its roof waiting at a glowing crosswalk, s
Share

Key points

  • Humanoid robots have advanced rapidly in movement and vision, but their ability to hear and communicate remains far behind.
  • The human brain devotes roughly 15% to 20% of its sensory processing to hearing, making it the second most energy-intensive sense after sight.
  • Real-world noise, echoes, and unpredictable acoustics cause current robot audio systems to fail in ways that quickly erode user trust.
  • High-quality audio training data is expensive and hard to collect, and simulating realistic sound requires far more computing power than simulating visuals.
  • Audio simulation company Treble is among a group of firms working to close the gap by building physically accurate sound models for training AI systems.

At a tech trade show, a humanoid robot walks across a crowded floor, picks up a cup, and hands it over without a wobble. Impressive. Then someone tries to give it a verbal instruction over the noise of the crowd. The robot mishears, pauses, and responds out of sequence. The spell is broken.

That awkward moment, repeated at event after event, is the focus of a new analysis published by The Robot Report, drawing on the perspective of Sigtryggur Kari Kristinsson, chief executive of Treble, a company that builds audio AI. His argument is straightforward: a robot that moves well but communicates badly will not be trusted, and a robot that is not trusted will not be adopted.

Why is robot hearing so far behind robot vision?

Vision came first because the data was easier. Cameras produce images that are simple to store, label, and feed into training systems. Audio is messier.

Sound changes depending on the shape of a room, the materials on the walls, and whether a source is moving. A voice recorded in a warehouse sounds completely different from the same voice in a kitchen. Collecting enough audio data to cover all those variations is slow and costly. Simulating it accurately requires solving complex wave-physics equations, which demands serious computing power.

Training platforms like NVIDIA Isaac Sim have helped robotics teams iterate quickly on movement and vision. Those environments are largely silent. That is a rational trade-off under real engineering constraints, but it means an entire layer of perception has been left underdeveloped.

What does this actually mean for people who use or work near robots?

A wobbly walk is forgivable. A robot that consistently mishears instructions, or replies at the wrong moment, feels broken even if everything else works perfectly.

Kristinsson frames this through evolutionary biology. Humans dedicate roughly 15% to 20% of sensory brain processing to hearing. Evolution is not wasteful. That allocation exists because sound tells us things vision cannot: whether something is approaching from behind, whether a voice is calm or distressed, whether an environment is safe. A robot that lacks that layer is, in practical terms, partly blind to the world it is supposed to operate in.

The implication for workers on a factory floor, patients in a care home, or anyone sharing a space with a humanoid robot is this: until robots can reliably hear and respond in noisy, real-world conditions, keep interactions simple and expect to repeat yourself.

What happens next?

Change is coming, slowly. Treble and a handful of other companies are building physically accurate sound-simulation tools, the audio equivalent of what visual simulators did for robot movement. The goal is to generate large, realistic audio training datasets without sending engineers into every possible real-world environment.

It is not a quick fix. Accurate acoustic simulation is computationally demanding, and the field is years behind its visual counterpart. But the direction is clear: robots that can hear properly will feel safer, more useful, and more trustworthy than those that cannot.

Common questions

Do any robots today handle noisy environments well?

Most commercial humanoid robots cope adequately in quiet, controlled settings, but performance drops significantly in loud or reverberant spaces like warehouses, kitchens, or crowded public areas.

Is this just a microphone problem?

No. Better microphones help, but the real challenge is software: training AI systems to separate a relevant voice from background noise, understand spatial cues, and respond in real time across wildly different acoustic conditions.

Will better audio AI make robots safer?

Yes, in practical terms. A robot that correctly hears a "stop" command or a warning shout in a noisy environment is meaningfully safer than one that cannot.

© 2026 AI2Day