AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Open ASR Leaderboard Adds Its First Global South Language on ThorstenMeyerAI.com

TL;DR

The Open ASR Leaderboard on Hugging Face has introduced Hindi and Indian English evaluation sets, making Hindi the first Indic and Global South language included. This expands the benchmark’s diversity and aims to improve speech recognition for underrepresented languages.

The Open ASR Leaderboard on Hugging Face has officially added two new evaluation sets, Monsoon en-IN for Indian English and Monsoon hi-IN for Hindi, marking the first inclusion of Indic and Global South languages on the platform. This development broadens the scope of the benchmark, which previously focused solely on European languages, and introduces a new level of diversity by including speech data from over 4,800 speakers across India. For more details, see the original analysis. The addition aims to improve speech recognition models’ performance on underrepresented languages and demographics, addressing a longstanding gap in benchmarking resources. Learn more about the significance of this development in the original analysis.

The new datasets comprise four splits: a public and private set for each language, with hours of audio and thousands of speakers. The Indian English sets feature 5.62 hours (public) and 5.58 hours (private) from over 2,800 speakers, collected from diverse regions using various devices and environments. The Hindi sets include 1.33 hours (public) from 468 speakers and 4.47 hours (private) from 1,571 speakers, gathered from spontaneous conversations across multiple districts. Each clip is annotated with 12 speaker attributes, including age, gender, occupation, education, income, and device type, to facilitate detailed analysis of model performance across different populations.

The datasets were designed to reflect real-world variability, capturing speech from individuals in different geographic locations, social backgrounds, and acoustic conditions. This approach aligns with efforts to improve speech recognition for underrepresented languages, as discussed in the original analysis. Contributors used their own devices and environments, ensuring the data’s authenticity. The datasets are split into public and private portions to prevent overfitting and gaming of benchmarks, with the private sets kept confidential to promote genuine model evaluation. The Hindi data employs a novel spelling lattice approach to account for spelling variations, unlike the normalisation used in English datasets.

At a glance
updateWhen: announced March 2024
The developmentThe Open ASR Leaderboard has added Hindi and Indian English datasets, representing the first Indic and Global South languages on the platform, with detailed speaker metadata and diverse audio conditions.
At a glance
announcementWhen: announced now; sets released publicly w…
The developmentVoice Arena and Hugging Face have added Hindi and Indian English evaluation sets — Monsoon hi-IN and Monsoon en-IN — to the Open ASR Leaderboard, making Hindi the first Global South language it covers.

Expanding Benchmark Diversity with Indic Languages

The inclusion of Hindi and Indian English on the Open ASR Leaderboard is a significant step toward diversifying speech recognition benchmarks, which have historically focused on European languages. Hindi, spoken by over half a billion people, has long been underrepresented in such benchmarks, and its addition signals a move toward more inclusive and representative evaluation standards. This development provides a new market signal to developers and researchers, encouraging the creation of models that perform well across different languages, dialects, and demographic groups. It also highlights the importance of detailed speaker metadata in understanding and mitigating biases in ASR systems, addressing disparities identified in prior research, such as racial and age-related performance gaps.

By enabling disaggregated analysis based on attributes like age, gender, and region, the new datasets aim to foster the development of more equitable speech recognition technologies. This progress could influence the future design of ASR models, pushing the industry toward fairness and inclusivity, especially for languages and populations historically marginalized in AI research.

Amazon

speech recognition microphone for Indian languages

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on the Open ASR Leaderboard and Dataset Expansion

The Open ASR Leaderboard, hosted on Hugging Face, has been a prominent benchmark for evaluating speech recognition models, primarily focusing on European languages such as English, French, and German. Its design emphasizes transparency and robustness by including both public and private splits, with a focus on preventing overfitting and model gaming. Recent research has shown that traditional metrics like Word Error Rate (WER) can mask disparities in model performance across different demographic groups, prompting efforts to incorporate more detailed speaker metadata and diverse datasets.

The platform’s expansion to include Hindi and Indian English follows a broader industry push to diversify benchmarks and address bias. Previous efforts in the field highlighted the need for datasets that reflect the linguistic and demographic diversity of global populations, especially in underrepresented regions like South Asia. The datasets are sourced from spontaneous, unscripted conversations, capturing natural speech in various acoustic environments, and are designed to challenge models with real-world variability.

This move aligns with ongoing initiatives to improve AI fairness and inclusivity, recognizing that speech recognition systems must serve diverse user bases effectively. The datasets’ emphasis on speaker attributes and geographic diversity aims to facilitate more nuanced evaluations and foster the development of models that perform reliably across different populations.

“Adding Hindi and Indian English datasets to the leaderboard marks a pivotal step toward more inclusive and representative speech recognition benchmarking.”

— Thorsten Meyer, AI researcher and contributor

Amazon

AI voice assistant for Hindi

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Dataset Impact and Model Performance

It remains unclear how existing top-performing models will perform on the new Hindi and Indian English datasets, as baseline results have not yet been published. The stability of model rankings given the relatively small size of the Hindi public split (1.33 hours) is also uncertain. Additionally, the effectiveness of the lattice approach for Hindi spelling variation against traditional normalisation methods has not been demonstrated through published comparisons. The impact of these datasets on the leaderboard’s overall evaluation metrics and whether participants will leverage the detailed speaker metadata for disaggregated reporting are still to be seen.

Amazon

language translation device Hindi English

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Benchmarking and Model Development

The datasets are now available for public and private testing, allowing researchers and developers to evaluate their models on Indian English and Hindi. Future work will likely include publishing baseline results on these new sets, assessing the stability of rankings, and exploring how the detailed speaker attributes influence model performance. Continued efforts may involve expanding the datasets further, increasing hours of audio, and encouraging the community to analyze results by demographic factors. Additionally, the platform may integrate these languages into broader benchmarks to promote inclusive AI development.

Amazon

Indian English speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is including Hindi on the Open ASR Leaderboard important?

Hindi is spoken by over half a billion people, yet it has been underrepresented in speech recognition benchmarks. Including it helps develop more inclusive models that serve a larger, more diverse population and addresses existing biases in AI systems.

How are the Hindi datasets different from previous benchmarks?

The Hindi datasets use a lattice approach to account for spelling variations, reflecting real-world differences more accurately than standard normalisation. They also include detailed speaker metadata and are sourced from spontaneous conversations across diverse environments.

Will current speech recognition models perform well on these new datasets?

Baseline results have not yet been published, so it is uncertain how existing models will perform. The datasets are designed to challenge models with real-world variability, which may reveal gaps in current systems.

What are the implications for AI fairness and bias reduction?

By providing detailed speaker attributes and diverse data, the datasets enable analysis of model performance across demographics, helping to identify and reduce biases in speech recognition technology.

What is the next step for the Open ASR Leaderboard with these languages?

Researchers and developers can now evaluate their models on the new datasets, publish results, and work toward improving model accuracy and fairness across Indian English and Hindi speakers.

Primary source: Hugging Face · via ThorstenMeyerAI.com

You May Also Like

Fully autonomous drones have killed human soldiers for the first time

Ukrainian officials confirm a test involving AI-controlled drones that killed soldiers on the battlefield, marking a significant development in autonomous warfare.

Anthropic Is Turning Claude Code’s Auto Mode On By Default – TechCrunch

Anthropic has made auto mode the default for Claude Code sessions on Pro, Max, and Team plans, enabling AI to perform actions with minimal human approval.

Two Decades Of RISC OS Open: A Reflection On Tech Operations Evolution

Celebrating 20 years of RISC OS Open, this article explores its development, impact, and ongoing significance in the tech landscape.

He made your free video player run smoothly. Now he’s doing that for robots.

Jean-Baptiste Kempf, creator of VLC Media Player, is now developing Kyber, a platform for real-time control of robots and drones, backed by $5M funding.