CoolFace
20 results

sarvam

sarvamai /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.audioautomatic-speech-recognition1K<n<10K18 likes1.4k downloads1mo agoHugging Facesarvamai /mmlu-indic Indic MMLU Dataset A multilingual version of the Massive Multitask Language Understanding (MMLU) benchmark, translated from English into 10 Indian languages. This version contains the translations of the development and test sets only. Languages Covered The dataset includes translations in the following languages: Bengali (bn) Gujarati (gu) Hindi (hi) Kannada (kn) Marathi (mr) Malayalam (ml) Oriya (or) Punjabi (pa) Tamil (ta) Telugu (te) Task Format Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/mmlu-indic.textquestion-answering100K<n<1M14 likes581 downloads1y agoHugging Facesarvamai /gsm8k-indictext10K<n<100K2 likes548 downloads1y agoHugging Facesarvamai /olmOCR-Bench-English olmOCR-bench (English Only) This is a filtered version of the allenai/olmOCR-bench dataset containing only English documents. Test Cases: Before vs After Category Before After Removed % Retained arxiv_math 2927 2917 10 99.7% headers_footers 760 520 240 68.4% long_tiny_text 442 442 0 100.0% multi_column 884 691 193 78.2% old_scans 526 526 0 100.0% old_scans_math 458 458 0 100.0% tables 1022 929 93 90.9% TOTAL 7019 6483 536 92.4% PDF Files:… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/olmOCR-Bench-English.documentimage-to-text1K<n<10K5 likes547 downloads8mo agoHugging Facesarvamai /samvaad-hi-v1100k high-quality conversations in English, Hindi, and Hinglish curated exclusively with an Indic context. texttext-generation100K<n<1M69 likes345 downloads2y agoHugging Faceritvik-sarvam /openswe-harbor OpenSWE-Harbor — NeMo Gym ready ⚠️ Read this before training: there are TWO sets in here — full (45,316) and filtered (8,875). Set Tasks Where it is When to use Full 45,316 routing/openswe_oss.jsonl, routing/openswe_other.jsonl, all of tasks/ Eval-only, dataset analysis, sweeps where you don't care about RL signal quality Filtered (RL default) 8,875 routing/openswe_oss_filtered.jsonl, routing/openswe_other_filtered.jsonl, filtered_ids.txt Use this for RL… See the full description on the dataset page: https://huggingface.co/datasets/ritvik-sarvam/openswe-harbor.text10K<n<100K0 likes278 downloads4mo agoHugging Face