Sarvam AI has launched Indic DiarBench, an open benchmark and dataset designed to improve artificial intelligence systems that can recognize speech and distinguish between different speakers across India’s 22 scheduled languages. The benchmark addresses a major gap in speech AI evaluation, where existing datasets often provide limited coverage of India’s linguistic diversity, dialects, code-mixed conversations and real-world multi-speaker environments.

The resource contains approximately 108 hours of natural multi-speaker audio, collected from near-field meetings, far-field recordings and real-world conversations. Its human-corrected annotations include time-aligned transcripts attributed to individual speakers, allowing researchers to evaluate both automatic speech recognition (ASR) and speaker diarization systems.

What Is Indic DiarBench?

Indic DiarBench is a multilingual benchmark focused on two closely related speech-AI tasks: automatic speech recognition and speaker diarization.

ASR systems convert spoken language into text, while speaker diarization determines who spoke when in a conversation.

The benchmark is designed to test both capabilities together in situations that more closely resemble real-world Indian speech.

Benchmark Snapshot

FeatureDetails
BenchmarkIndic DiarBench
Language CoverageAll 22 scheduled Indian languages
AudioApproximately 108 hours
Main TasksASR and speaker diarization
Audio TypesMeetings, far-field and in-the-wild recordings
AnnotationsHuman-corrected, time-aligned speaker transcripts
AccessOpen-access

Designed for India’s Linguistic Diversity

Indian speech presents challenges that are not always captured by conventional speech benchmarks.

Indic DiarBench incorporates several characteristics common in everyday conversations, including:

  • Regional dialect variations.
  • English code-mixing.
  • Multiple people speaking simultaneously.
  • Different recording environments.
  • Natural conversational patterns.
  • Far-field speech.

These characteristics make the benchmark more representative of the conditions AI systems encounter outside controlled laboratory environments.

Why Speaker Diarization Matters

Speaker diarization is particularly important for applications that need to understand conversations involving multiple people.

Potential applications include:

  • Call-centre transcription.
  • Meeting summaries.
  • Government service interactions.
  • Healthcare conversations.
  • Legal transcription.
  • Customer-service analytics.
  • Voice assistants.
  • Multilingual enterprise applications.

For example, simply converting a conversation into text is not enough when several people are speaking. An effective system must also determine which words belong to which participant.

Benchmark Covers All 22 Scheduled Languages

One of the most significant aspects of Indic DiarBench is its coverage of all 22 scheduled languages of India.

This is important because many existing AI benchmarks are heavily concentrated around English and a smaller number of globally prominent languages.

A benchmark covering India’s full scheduled-language ecosystem can help researchers identify where speech models perform well and where they struggle.

Key Challenges Tested

ChallengeWhy It Matters
Code-MixingIndian speakers frequently combine English with regional languages
DialectsPronunciation and vocabulary vary across regions
Speaker OverlapPeople often interrupt or speak simultaneously
Far-Field AudioReal conversations may be recorded from a distance
Natural SpeechEveryday speech differs from scripted recordings

Human-Corrected Data

Indic DiarBench uses human-corrected annotations, including time-aligned transcripts linked to individual speakers.

This provides researchers with a reliable reference against which AI systems can be evaluated.

High-quality annotation is particularly important for speaker diarization because an AI system must correctly identify both the spoken words and the timing and identity of each speaker.

Testing Current AI Systems

The researchers also evaluated leading speech APIs and multimodal large language models using Indic DiarBench.

The goal is not simply to publish another dataset but to establish a common testing environment for comparing how well modern AI systems handle Indian multilingual conversations.

The benchmark can therefore help researchers measure progress over time and identify specific weaknesses in existing speech technologies.

Building India’s Speech AI Ecosystem

Indic DiarBench fits into Sarvam AI’s broader focus on developing AI technology specifically for Indian languages.

The company has developed models and tools covering speech recognition, text-to-speech, translation, document intelligence and conversational AI. Sarvam says its platform supports India’s 22 languages and is designed for enterprise, government and developer applications.

The company has also been expanding its broader AI model portfolio, including multilingual large language models and speech technologies designed for Indian users.

Why Open Benchmarks Matter

Open benchmarks can play an important role in improving AI because they allow researchers, startups and technology companies to evaluate systems using the same reference data.

For Indian-language AI, this could help:

  • Identify performance gaps between languages.
  • Improve speech recognition accuracy.
  • Develop better multilingual models.
  • Encourage independent research.
  • Create more reliable voice-based applications.

It could also make it easier to compare commercial and open-source systems on realistic Indian-language speech rather than relying primarily on English-centric evaluations.

Looking Ahead

Sarvam AI’s Indic DiarBench provides researchers with a dedicated testing ground for one of the most challenging areas of multilingual AI: understanding natural conversations involving multiple speakers across India’s diverse languages. With around 108 hours of human-annotated audio spanning all 22 scheduled languages, the benchmark is designed to capture real-world factors such as code-mixing, dialect variation, overlapping speech and different recording environments.

Looking ahead, wider adoption of open benchmarks such as Indic DiarBench could help accelerate the development of more accurate Indian-language voice technologies. As voice AI expands across government services, customer support, healthcare, education and enterprise applications, better evaluation tools will become increasingly important for ensuring that AI systems work reliably across India’s linguistic and cultural diversity.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.