datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.mmlu-indic
Indic MMLU Dataset
A multilingual version of the Massive Multitask Language Understanding (MMLU) benchmark, translated from English into 10 Indian languages.
This version contains the translations of the development and test sets only.
Languages Covered
The dataset includes translations in the following languages:
Bengali (bn)
Gujarati (gu)
Hindi (hi)
Kannada (kn)
Marathi (mr)
Malayalam (ml)
Oriya (or)
Punjabi (pa)
Tamil (ta)
Telugu (te)
Task Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/mmlu-indic.gsm8k-indicsamvaad-hi-v1100k high-quality conversations in English, Hindi, and Hinglish curated exclusively with an Indic context.
openswe-harbor
OpenSWE-Harbor — NeMo Gym ready
⚠️ Read this before training: there are TWO sets in here — full (45,316) and filtered (8,875).
Set
Tasks
Where it is
When to use
Full
45,316
routing/openswe_oss.jsonl, routing/openswe_other.jsonl, all of tasks/
Eval-only, dataset analysis, sweeps where you don't care about RL signal quality
Filtered (RL default)
8,875
routing/openswe_oss_filtered.jsonl, routing/openswe_other_filtered.jsonl, filtered_ids.txt
Use this for RL… See the full description on the dataset page: https://huggingface.co/datasets/ritvik-sarvam/openswe-harbor.tatoeba-indic
Tatoeba Benchmark (Indian languages only)
This benchmark is prepared from the 2023 Tatoeba Challenge, by extracting the dev and test sets for languages spoken in the Indian Republic.
The code to download and process the data can be found here in the repo: data_prep/original_v1/extract.py
Note: This is not the official version of Tatoeba benchmark. Just a processed mirror for Indian languages, made available in HuggingFace for ease of use.
Languages
Language code… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tatoeba-indic.vagartha
वागर्थ · Vāgartha
वागर्थाविव संपृक्तौ वागर्थप्रतिपत्तये ।जगतः पितरौ वन्दे पार्वतीपरमेश्वरौ ॥
"United as word and meaning are united, I bow to the parents of the world,Pārvatī and Parameśvara, that I may attain an understanding of word and meaning."
— Kālidāsa, Raghuvaṃśa 1.1
Vāgartha — vāk (word) and artha (meaning) — is a corpus of 217,959 Sanskrit
verses, each paired with a detailed, structured explanation in English. The name is
taken from the invocation above, in which… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/vagartha.arc-challenge-indicboolq-indic
Indic BoolQ Dataset
A multilingual version of the BoolQ (Boolean Questions) dataset, translated from English into 10 Indian languages.
It is a question-answering dataset for yes/no questions containing ~12k naturally occurring questions.
Languages Covered
The dataset includes translations in the following languages:
Bengali (bn)
Gujarati (gu)
Hindi (hi)
Kannada (kn)
Marathi (mr)
Malayalam (ml)
Oriya (or)
Punjabi (pa)
Tamil (ta)
Telugu (te)
Dataset Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/boolq-indic.sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600Dataset for sarvam's entity normalisation task. More detailed information can be found here, in the main model repo: Hugging Face
Detailed Report (Writeup): Google Drive
It also has a gguf variant, with certain additional gguf based innstructions: Hugging Face
Model inference script can be found here: Colab
Model predictions can be found in this dataset and both the repo files. named as:
eval_data_001_predictions.csv and eval_data_001_predictions_excel.csv.
train_data_001_predictions.csvand… See the full description on the dataset page: https://huggingface.co/datasets/Tasmay-Tib/sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600.indic-safety-eval
IndicSafetyBench
Multi-turn safety evaluation benchmark for Indian languages. Tests models against 30 jailbreak techniques across 19 India-specific harm domains in 23 languages with 189 dialect varieties.
Stats
3,374 benchmark items
30 jailbreak techniques across 8 families
19 India harm domains (territorial, caste, religious, political, gender, etc.)
14 global harm categories (aligned with AILuminate v1.0 / Llama Guard 4)
23 languages (22 Indic + English)
~29%… See the full description on the dataset page: https://huggingface.co/datasets/anna-sarvam/indic-safety-eval.contextual_asr_benchmark
Synthetic Contextual ASR Benchmark (Indic)
Dataset Summary
This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses.
The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/contextual_asr_benchmark.indivibe
indivibe
Overview
To rigorously evaluate the Indic capabilities of Sarvam models across all 22 scheduled languages, we developed a new Indic benchmark and evaluated models using a pairwise comparison framework with an LLM-as-judge protocol. A key goal of this benchmark is to reflect how language is actually used in India today. In practice, this means evaluating each language in two script styles: native script (formal written usage) and romanized Latin script (colloquial… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indivibe.trivia-qa-indic-mcqtts-general-benchmark
TTS General Benchmark
A multilingual Text-to-Speech (TTS) evaluation benchmark covering 11 Indian languages across multiple real-world use cases. The dataset is designed for systematic and repeatable evaluation of TTS systems under both high-quality and telephony bandwidth conditions.
Total prompts: 1,815 unique text samplesLanguages: 11Evaluation tracks: High Quality + 8 kHz Telephony
This is an evaluation-only benchmark dataset intended for testing and comparison — not for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tts-general-benchmark.indic-ocr-bench
Sarvam Indic OCR Bench
Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material.
The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.Sarvam-105b-Distill-100k
Sarvam 105B Distill 100K
Dataset Summary
Reasoning dataset distilled from Sarvam 105B packaged with / tags in thinking/sharegpt/chatml/simple_qa schemas.
Source
Input JSONL: distillation_pipeline\dataset_final_p1_100k\full_dataset.jsonl
Generated at: 2026-05-17T11:24:23.743584+00:00
Splits
Train: 92040
Validation: 1917
Test: 1918
Distribution Counts
Domain
coding_computer_science: 16667
creative_planning_openended: 3809… See the full description on the dataset page: https://huggingface.co/datasets/abhinav0231/Sarvam-105b-Distill-100k.tts-robustness-benchmark
TTS Robustness Benchmark
Overview
The TTS Robustness Benchmark is a set of evaluation samples used to calculate domain-wise CER (Character Error Rate) scores across 7 critical stress-test categories. This benchmark ensures that each sentence-language pair appears exactly once, providing a clean and reliable metric for TTS robustness.
Blog Post: Read the official announcement
Column Descriptions
Column
Description
text
The original input… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tts-robustness-benchmark.google-dakshinasarvam-tts-in-te-en
Indian English + Telugu Single-Speaker TTS Dataset (emotion-tagged)
Clean audio clips sourced from YouTube, transcribed with Sarvam ASR, segmented with
diarization, and labeled with emotion/style tags. Built as a data-quality / curation exercise.
"Single-speaker" means each clip contains exactly one speaker (verified by
diarization and speaker-embedding similarity). The dataset spans 9 distinct speakers
total (4 English, 5 Telugu), tracked via speaker_id.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/AkCodes23/sarvam-tts-in-te-en.audiollm-evalsThis evaluation set contains ~100 questions in both text and audio format in Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu. We use this dataset internally at Sarvam to evaluate the performance of our audio models. We open-source this data to enable the research community to replicate the results mentioned in our Shuka blog.
By deisgn, the questions are sometimes vague, and the audio has noise and other inconsistencies, to measure the robustness of… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/audiollm-evals.sarvam-tts-dataset
Sarvam-style TTS Training Dataset (Telugu + Indian English)
A small, high-curation single-speaker-per-clip TTS dataset (~80 minutes, 193 clips) of
Telugu and Indian English speech. Every clip ships an accurate transcript and a
rich, Parler-TTS-style natural-language style/emotion description. Built as a take-home
where the emphasis is data quality and curation judgment, not pipeline code — "listen
to the data" was the core principle.
Clips: 193 (~25s each, always cut at silence… See the full description on the dataset page: https://huggingface.co/datasets/theshaikasad/sarvam-tts-dataset.sarvam-dub-benchmark-set
Sarvam Dubbing Benchmark Dataset
Dataset Description
Multilingual evaluation dataset for real-time dubbing and voice cloning benchmarking with focus on speaker similarity preservation across same-lingual and cross-lingual scenarios.
This dataset was used to benchmark production dubbing systems. Internal evaluations showed higher speaker similarity than ElevenLabs v3 and Cartesia Sonic under an identical scoring protocol.
Check out the Sarvam Dub blog for more… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/sarvam-dub-benchmark-set.chat_data_v1sarvam-indian-eng-hin-tts
Indian English + Hindi TTS Dataset (emotion-tagged)
A curated, single-speaker-per-clip speech dataset for Text-to-Speech research,
covering Indian English and Hindi. Every clip is sourced from YouTube,
transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM.
Total: 82 clips, 55.6 minutes
Hindi: 28.8 min | Indian English: 26.8 min
Audio: mono, 24 kHz, 16-bit WAV
Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.helpsteer-dposft_saaras_flow
sft_saaras_flow
Synthetic SFT data for ASR transcription correction, generated to distill a large
model (Gemini 3.1 Pro) into a smaller correction model. Each example is a realistic raw
ASR transcript (in the spoken language's native script) plus the app configuration
(variation axes) and the corrected output the app's real formatting prompt produces.
1100 examples, 11 languages.
Raw transcripts are in the native script of the spoken language; mode controls the
correction… See the full description on the dataset page: https://huggingface.co/datasets/sarvam/sft_saaras_flow.distilled-data-sarvam-105bsarvam-tts-dataset
Sarvam Indian TTS Dataset
A curated Text-to-Speech dataset containing Indian English and Hindi
single-speaker audio segments, built as part of the Sarvam AI hiring assignment.
Dataset Summary
Property
Value
Total duration
78.8 minutes
English (en-IN)
57.5 minutes
Hindi (hi-IN)
21.3 minutes
Total segments
175
Human reviewed
99.4%
Mean quality score
92.1/100
Mean SNR
49.0 dB
Mean ASR confidence
0.900
Dataset Card… See the full description on the dataset page: https://huggingface.co/datasets/Avishii0309/sarvam-tts-dataset.nvidia_math_r1_llama
