CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sarvamai /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.audioautomatic-speech-recognition1K<n<10K18 likes1.4k downloads2mo agoHugging Face02sarvamai /mmlu-indic Indic MMLU Dataset A multilingual version of the Massive Multitask Language Understanding (MMLU) benchmark, translated from English into 10 Indian languages. This version contains the translations of the development and test sets only. Languages Covered The dataset includes translations in the following languages: Bengali (bn) Gujarati (gu) Hindi (hi) Kannada (kn) Marathi (mr) Malayalam (ml) Oriya (or) Punjabi (pa) Tamil (ta) Telugu (te) Task Format Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/mmlu-indic.textquestion-answering100K<n<1M14 likes581 downloads1y agoHugging Face03sarvamai /gsm8k-indictext10K<n<100K2 likes548 downloads1y agoHugging Face04sarvamai /samvaad-hi-v1100k high-quality conversations in English, Hindi, and Hinglish curated exclusively with an Indic context. texttext-generation100K<n<1M69 likes345 downloads2y agoHugging Face05ritvik-sarvam /openswe-harbor OpenSWE-Harbor — NeMo Gym ready ⚠️ Read this before training: there are TWO sets in here — full (45,316) and filtered (8,875). Set Tasks Where it is When to use Full 45,316 routing/openswe_oss.jsonl, routing/openswe_other.jsonl, all of tasks/ Eval-only, dataset analysis, sweeps where you don't care about RL signal quality Filtered (RL default) 8,875 routing/openswe_oss_filtered.jsonl, routing/openswe_other_filtered.jsonl, filtered_ids.txt Use this for RL… See the full description on the dataset page: https://huggingface.co/datasets/ritvik-sarvam/openswe-harbor.text10K<n<100K0 likes278 downloads4mo agoHugging Face06sarvamai /tatoeba-indic Tatoeba Benchmark (Indian languages only) This benchmark is prepared from the 2023 Tatoeba Challenge, by extracting the dev and test sets for languages spoken in the Indian Republic. The code to download and process the data can be found here in the repo: data_prep/original_v1/extract.py Note: This is not the official version of Tatoeba benchmark. Just a processed mirror for Indian languages, made available in HuggingFace for ease of use. Languages Language code… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tatoeba-indic.text10K<n<100K2 likes250 downloads1y agoHugging Face07sarvamai /vagartha वागर्थ · Vāgartha वागर्थाविव संपृक्तौ वागर्थप्रतिपत्तये ।जगतः पितरौ वन्दे पार्वतीपरमेश्वरौ ॥ "United as word and meaning are united, I bow to the parents of the world,Pārvatī and Parameśvara, that I may attain an understanding of word and meaning." — Kālidāsa, Raghuvaṃśa 1.1 Vāgartha — vāk (word) and artha (meaning) — is a corpus of 217,959 Sanskrit verses, each paired with a detailed, structured explanation in English. The name is taken from the invocation above, in which… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/vagartha.texttext-generation100K<n<1M10 likes161 downloads21d agoHugging Face08sarvamai /arc-challenge-indictext10K<n<100K3 likes151 downloads1y agoHugging Face09sarvamai /boolq-indic Indic BoolQ Dataset A multilingual version of the BoolQ (Boolean Questions) dataset, translated from English into 10 Indian languages. It is a question-answering dataset for yes/no questions containing ~12k naturally occurring questions. Languages Covered The dataset includes translations in the following languages: Bengali (bn) Gujarati (gu) Hindi (hi) Kannada (kn) Marathi (mr) Malayalam (ml) Oriya (or) Punjabi (pa) Tamil (ta) Telugu (te) Dataset Format Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/boolq-indic.textquestion-answering100K<n<1M0 likes102 downloads2y agoHugging Face10Tasmay-Tib /sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600Dataset for sarvam's entity normalisation task. More detailed information can be found here, in the main model repo: Hugging Face Detailed Report (Writeup): Google Drive It also has a gguf variant, with certain additional gguf based innstructions: Hugging Face Model inference script can be found here: Colab Model predictions can be found in this dataset and both the repo files. named as: eval_data_001_predictions.csv and eval_data_001_predictions_excel.csv. train_data_001_predictions.csvand… See the full description on the dataset page: https://huggingface.co/datasets/Tasmay-Tib/sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600.text1K<n<10K0 likes101 downloads2y agoHugging Face11anna-sarvam /indic-safety-eval IndicSafetyBench Multi-turn safety evaluation benchmark for Indian languages. Tests models against 30 jailbreak techniques across 19 India-specific harm domains in 23 languages with 189 dialect varieties. Stats 3,374 benchmark items 30 jailbreak techniques across 8 families 19 India harm domains (territorial, caste, religious, political, gender, etc.) 14 global harm categories (aligned with AILuminate v1.0 / Llama Guard 4) 23 languages (22 Indic + English) ~29%… See the full description on the dataset page: https://huggingface.co/datasets/anna-sarvam/indic-safety-eval.texttext-generation1K<n<10K0 likes94 downloads3mo agoHugging Face12sarvamai /contextual_asr_benchmark Synthetic Contextual ASR Benchmark (Indic) Dataset Summary This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses. The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/contextual_asr_benchmark.audio1K<n<10K4 likes93 downloads8mo agoHugging Face13sarvamai /indivibe indivibe Overview To rigorously evaluate the Indic capabilities of Sarvam models across all 22 scheduled languages, we developed a new Indic benchmark and evaluated models using a pairwise comparison framework with an LLM-as-judge protocol. A key goal of this benchmark is to reflect how language is actually used in India today. In practice, this means evaluating each language in two script styles: native script (formal written usage) and romanized Latin script (colloquial… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indivibe.texttext-generation1K<n<10K2 likes84 downloads5mo agoHugging Face14sarvamai /trivia-qa-indic-mcqtext100K<n<1M3 likes82 downloads1y agoHugging Face15sarvamai /tts-general-benchmark TTS General Benchmark A multilingual Text-to-Speech (TTS) evaluation benchmark covering 11 Indian languages across multiple real-world use cases. The dataset is designed for systematic and repeatable evaluation of TTS systems under both high-quality and telephony bandwidth conditions. Total prompts: 1,815 unique text samplesLanguages: 11Evaluation tracks: High Quality + 8 kHz Telephony This is an evaluation-only benchmark dataset intended for testing and comparison — not for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tts-general-benchmark.texttext-to-speech1K<n<10K7 likes79 downloads8mo agoHugging Face16sarvamai /indic-ocr-bench Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material. The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.imageimage-to-text1K<n<10K6 likes57 downloads4d agoHugging Face17abhinav0231 /Sarvam-105b-Distill-100k Sarvam 105B Distill 100K Dataset Summary Reasoning dataset distilled from Sarvam 105B packaged with / tags in thinking/sharegpt/chatml/simple_qa schemas. Source Input JSONL: distillation_pipeline\dataset_final_p1_100k\full_dataset.jsonl Generated at: 2026-05-17T11:24:23.743584+00:00 Splits Train: 92040 Validation: 1917 Test: 1918 Distribution Counts Domain coding_computer_science: 16667 creative_planning_openended: 3809… See the full description on the dataset page: https://huggingface.co/datasets/abhinav0231/Sarvam-105b-Distill-100k.texttext-generation100K<n<1M1 likes48 downloads4mo agoHugging Face18sarvamai /tts-robustness-benchmark TTS Robustness Benchmark Overview The TTS Robustness Benchmark is a set of evaluation samples used to calculate domain-wise CER (Character Error Rate) scores across 7 critical stress-test categories. This benchmark ensures that each sentence-language pair appears exactly once, providing a clean and reliable metric for TTS robustness. Blog Post: Read the official announcement Column Descriptions Column Description text The original input… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tts-robustness-benchmark.texttext-to-speechn<1K6 likes44 downloads8mo agoHugging Face19suhani-sarvam /google-dakshinatext100K<n<1M2 likes41 downloads2y agoHugging Face20AkCodes23 /sarvam-tts-in-te-en Indian English + Telugu Single-Speaker TTS Dataset (emotion-tagged) Clean audio clips sourced from YouTube, transcribed with Sarvam ASR, segmented with diarization, and labeled with emotion/style tags. Built as a data-quality / curation exercise. "Single-speaker" means each clip contains exactly one speaker (verified by diarization and speaker-embedding similarity). The dataset spans 9 distinct speakers total (4 English, 5 Telugu), tracked via speaker_id. Contents… See the full description on the dataset page: https://huggingface.co/datasets/AkCodes23/sarvam-tts-in-te-en.audiotext-to-speechn<1K0 likes38 downloads3mo agoHugging Face21sarvamai /audiollm-evalsThis evaluation set contains ~100 questions in both text and audio format in Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu. We use this dataset internally at Sarvam to evaluate the performance of our audio models. We open-source this data to enable the research community to replicate the results mentioned in our Shuka blog. By deisgn, the questions are sometimes vague, and the audio has noise and other inconsistencies, to measure the robustness of… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/audiollm-evals.audion<1K3 likes37 downloads2y agoHugging Face22theshaikasad /sarvam-tts-dataset Sarvam-style TTS Training Dataset (Telugu + Indian English) A small, high-curation single-speaker-per-clip TTS dataset (~80 minutes, 193 clips) of Telugu and Indian English speech. Every clip ships an accurate transcript and a rich, Parler-TTS-style natural-language style/emotion description. Built as a take-home where the emphasis is data quality and curation judgment, not pipeline code — "listen to the data" was the core principle. Clips: 193 (~25s each, always cut at silence… See the full description on the dataset page: https://huggingface.co/datasets/theshaikasad/sarvam-tts-dataset.audiotext-to-speechn<1K0 likes32 downloads3mo agoHugging Face23sarvamai /sarvam-dub-benchmark-set Sarvam Dubbing Benchmark Dataset Dataset Description Multilingual evaluation dataset for real-time dubbing and voice cloning benchmarking with focus on speaker similarity preservation across same-lingual and cross-lingual scenarios. This dataset was used to benchmark production dubbing systems. Internal evaluations showed higher speaker similarity than ElevenLabs v3 and Cartesia Sonic under an identical scoring protocol. Check out the Sarvam Dub blog for more… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/sarvam-dub-benchmark-set.audiotext-to-speechn<1K2 likes31 downloads8mo agoHugging Face24anna-sarvam /chat_data_v1text100K<n<1M0 likes31 downloads4mo agoHugging Face25Abhi29112005 /sarvam-indian-eng-hin-tts Indian English + Hindi TTS Dataset (emotion-tagged) A curated, single-speaker-per-clip speech dataset for Text-to-Speech research, covering Indian English and Hindi. Every clip is sourced from YouTube, transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM. Total: 82 clips, 55.6 minutes Hindi: 28.8 min &nbsp;|&nbsp; Indian English: 26.8 min Audio: mono, 24 kHz, 16-bit WAV Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.audiotext-to-speechn<1K0 likes31 downloads3mo agoHugging Face26tanay-sarvam /helpsteer-dpotext10K<n<100K0 likes29 downloads2y agoHugging Face27sarvam /sft_saaras_flow sft_saaras_flow Synthetic SFT data for ASR transcription correction, generated to distill a large model (Gemini 3.1 Pro) into a smaller correction model. Each example is a realistic raw ASR transcript (in the spoken language's native script) plus the app configuration (variation axes) and the corrected output the app's real formatting prompt produces. 1100 examples, 11 languages. Raw transcripts are in the native script of the spoken language; mode controls the correction… See the full description on the dataset page: https://huggingface.co/datasets/sarvam/sft_saaras_flow.texttext-generation1K<n<10K0 likes29 downloads3mo agoHugging Face28neuralnets /distilled-data-sarvam-105btext10K<n<100K0 likes28 downloads7mo agoHugging Face29Avishii0309 /sarvam-tts-dataset Sarvam Indian TTS Dataset A curated Text-to-Speech dataset containing Indian English and Hindi single-speaker audio segments, built as part of the Sarvam AI hiring assignment. Dataset Summary Property Value Total duration 78.8 minutes English (en-IN) 57.5 minutes Hindi (hi-IN) 21.3 minutes Total segments 175 Human reviewed 99.4% Mean quality score 92.1/100 Mean SNR 49.0 dB Mean ASR confidence 0.900 Dataset Card… See the full description on the dataset page: https://huggingface.co/datasets/Avishii0309/sarvam-tts-dataset.audiotext-to-speechn<1K0 likes28 downloads3mo agoHugging Face30aashay-sarvam /nvidia_math_r1_llamatext10K<n<100K3 likes26 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.