Speech AI Gets Its First Major Hindi Benchmark, and It Measures Who Gets Left Behind
A new evaluation dataset called Monsoon puts Hindi and Indian English on the global leaderboard for the first time, with nearly 5,000 speakers and data designed to expose the gaps that single-score tests hide.

Key points
- The Open ASR Leaderboard, a public ranking of AI speech-recognition tools, added its first South Asian languages in 2024: Hindi and Indian English.
- The Monsoon dataset covers 4,888 speakers across hundreds of Indian districts, making it one of the most geographically spread speech benchmarks ever built.
- Commercial speech-recognition systems have been shown to make roughly twice as many errors for Black speakers as for white speakers, according to published research.
- Hindi is spoken by more than half a billion people but was absent from every major multilingual speech leaderboard until now.
- Each Monsoon speaker has 12 pieces of recorded background information, including age, gender, region and device, so researchers can trace exactly where a model fails.
When an AI company wants to know whether its speech-recognition tool is any good, it runs the tool against a benchmark: a curated collection of recorded speech, with correct transcriptions already written out. The tool listens, types what it hears, and the score is how often it gets the words right. That score is called WER, short for word error rate, and a lower number means fewer mistakes.
The problem is that one number hides a lot.
A widely cited research paper found that major commercial systems made roughly twice as many errors on speech from Black Americans as on speech from white Americans. Separate work found further gaps by age, gender and accent. None of that shows up in the headline figure, because the recordings used in most benchmarks carry almost no information about the person speaking.
Why does this matter for ordinary people?
If a model scores well on a benchmark, it gets adopted. Doctors' dictation software, call-centre assistants, accessibility tools for people who cannot type: all of them depend on speech recognition, and all of them inherit whatever blind spots the benchmark missed.
Hindi speakers number more than half a billion people worldwide. Until now, every multilingual speech leaderboard covered only European languages. That meant AI developers had no standard way to measure how well their tools actually worked for one of the planet's largest speaking populations.
The Monsoon dataset, built through a partnership between Voice Arena and Hugging Face (the open-source AI platform that hosts the leaderboard), changes that. It adds two new test sets: one for Hindi, one for Indian English.
What makes Monsoon different from other benchmarks?
Most benchmark audio comes from whatever recordings were easy to find, often collected in quiet rooms on standardised microphones. Monsoon was built the opposite way.
Contributors used their own phones, in their own homes and streets. Recordings came from across India: the public Indian English set alone draws on 428 districts across 30 states. No single device model accounts for more than 2.1% of clips. The ten speakers who contributed the most audio together represent under 7% of the total, so no single voice can drag the score.
The dataset also records 12 background details per speaker, including age, gender, region, occupation, education and income band. That means researchers can ask not just "how did the model score?" but "who did it fail?"
Hindi posed a specific extra challenge. The language has many accepted spellings for the same word, and no standard automatic correction can resolve them. Monsoon handles this with a lattice (a list of every accepted spelling for each word or phrase in the transcript) so a correct answer is not accidentally penalised for using a different but valid spelling.
What should readers watch for?
If you use voice-to-text tools in Hindi or Indian-accented English, for note-taking, accessibility or work, pay attention to error patterns. If the tool struggles more when you speak in your natural accent than when you adjust toward a neutral accent, that gap is exactly what Monsoon is designed to surface and, eventually, fix.
Researchers and journalists tracking AI fairness should watch whether model developers begin reporting results broken down by speaker background, not just overall scores. The data now exists to demand that.



