Sarvam AI has launched Indic DiarBench, an open benchmark and dataset designed to improve artificial intelligence systems that can recognize speech and distinguish between different speakers across India’s 22 scheduled languages. The benchmark addresses a major gap in speech AI evaluation, where existing datasets often provide limited coverage of India’s linguistic diversity, dialects, code-mixed conversations and real-world multi-speaker environments.
The resource contains approximately 108 hours of natural multi-speaker audio, collected from near-field meetings, far-field recordings and real-world conversations. Its human-corrected annotations include time-aligned transcripts attributed to individual speakers, allowing researchers to evaluate both automatic speech recognition (ASR) and speaker diarization systems.
What Is Indic DiarBench?
Indic DiarBench is a multilingual benchmark focused on two closely related speech-AI tasks: automatic speech recognition and speaker diarization.
ASR systems convert spoken language into text, while speaker diarization determines who spoke when in a conversation.
The benchmark is designed to test both capabilities together in situations that more closely resemble real-world Indian speech.
Benchmark Snapshot
| Feature | Details |
|---|---|
| Benchmark | Indic DiarBench |
| Language Coverage | All 22 scheduled Indian languages |
| Audio | Approximately 108 hours |
| Main Tasks | ASR and speaker diarization |
| Audio Types | Meetings, far-field and in-the-wild recordings |
| Annotations | Human-corrected, time-aligned speaker transcripts |
| Access | Open-access |
Designed for India’s Linguistic Diversity
Indian speech presents challenges that are not always captured by conventional speech benchmarks.
Indic DiarBench incorporates several characteristics common in everyday conversations, including:
- Regional dialect variations.
- English code-mixing.
- Multiple people speaking simultaneously.
- Different recording environments.
- Natural conversational patterns.
- Far-field speech.
These characteristics make the benchmark more representative of the conditions AI systems encounter outside controlled laboratory environments.
Why Speaker Diarization Matters
Speaker diarization is particularly important for applications that need to understand conversations involving multiple people.
Potential applications include:
- Call-centre transcription.
- Meeting summaries.
- Government service interactions.
- Healthcare conversations.
- Legal transcription.
- Customer-service analytics.
- Voice assistants.
- Multilingual enterprise applications.
For example, simply converting a conversation into text is not enough when several people are speaking. An effective system must also determine which words belong to which participant.
Benchmark Covers All 22 Scheduled Languages
One of the most significant aspects of Indic DiarBench is its coverage of all 22 scheduled languages of India.
This is important because many existing AI benchmarks are heavily concentrated around English and a smaller number of globally prominent languages.
A benchmark covering India’s full scheduled-language ecosystem can help researchers identify where speech models perform well and where they struggle.
Key Challenges Tested
| Challenge | Why It Matters |
|---|---|
| Code-Mixing | Indian speakers frequently combine English with regional languages |
| Dialects | Pronunciation and vocabulary vary across regions |
| Speaker Overlap | People often interrupt or speak simultaneously |
| Far-Field Audio | Real conversations may be recorded from a distance |
| Natural Speech | Everyday speech differs from scripted recordings |
Human-Corrected Data
Indic DiarBench uses human-corrected annotations, including time-aligned transcripts linked to individual speakers.
This provides researchers with a reliable reference against which AI systems can be evaluated.
High-quality annotation is particularly important for speaker diarization because an AI system must correctly identify both the spoken words and the timing and identity of each speaker.
Testing Current AI Systems
The researchers also evaluated leading speech APIs and multimodal large language models using Indic DiarBench.
The goal is not simply to publish another dataset but to establish a common testing environment for comparing how well modern AI systems handle Indian multilingual conversations.
The benchmark can therefore help researchers measure progress over time and identify specific weaknesses in existing speech technologies.
Building India’s Speech AI Ecosystem
Indic DiarBench fits into Sarvam AI’s broader focus on developing AI technology specifically for Indian languages.
The company has developed models and tools covering speech recognition, text-to-speech, translation, document intelligence and conversational AI. Sarvam says its platform supports India’s 22 languages and is designed for enterprise, government and developer applications.
The company has also been expanding its broader AI model portfolio, including multilingual large language models and speech technologies designed for Indian users.
Why Open Benchmarks Matter
Open benchmarks can play an important role in improving AI because they allow researchers, startups and technology companies to evaluate systems using the same reference data.
For Indian-language AI, this could help:
- Identify performance gaps between languages.
- Improve speech recognition accuracy.
- Develop better multilingual models.
- Encourage independent research.
- Create more reliable voice-based applications.
It could also make it easier to compare commercial and open-source systems on realistic Indian-language speech rather than relying primarily on English-centric evaluations.
Looking Ahead
Sarvam AI’s Indic DiarBench provides researchers with a dedicated testing ground for one of the most challenging areas of multilingual AI: understanding natural conversations involving multiple speakers across India’s diverse languages. With around 108 hours of human-annotated audio spanning all 22 scheduled languages, the benchmark is designed to capture real-world factors such as code-mixing, dialect variation, overlapping speech and different recording environments.
Looking ahead, wider adoption of open benchmarks such as Indic DiarBench could help accelerate the development of more accurate Indian-language voice technologies. As voice AI expands across government services, customer support, healthcare, education and enterprise applications, better evaluation tools will become increasingly important for ensuring that AI systems work reliably across India’s linguistic and cultural diversity.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.

