
Sarvam AI has introduced 'Indic DiarBench', a new benchmark dataset designed to evaluate speech recognition and speaker attribution across 22 Indian languages.
The dataset provides two hours of annotated audio for each language, aiming to bridge the gap in AI performance for regional dialects and multilingual speech processing.
This initiative, developed in collaboration with AI4Bharat, is expected to accelerate the development of more accurate voice-based AI applications for the Indian market.