
Sarvam AI and AI4Bharat introduced Indic DiarBench, an open benchmark dataset covering all 22 scheduled Indian languages to evaluate speech recognition and speaker diarization.
The dataset includes 108 hours of natural, spontaneous speech from 485 speakers across 189 districts, capturing real-world challenges like overlapping speech and background noise.
This tool aims to improve the consistency of speech AI systems used in meeting transcriptions, customer service automation, and voice assistants across India's diverse linguistic landscape.