sarvam
Datasets
All datasets matching “sarvam”indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.mmlu-indic
Indic MMLU Dataset
A multilingual version of the Massive Multitask Language Understanding (MMLU) benchmark, translated from English into 10 Indian languages.
This version contains the translations of the development and test sets only.
Languages Covered
The dataset includes translations in the following languages:
Bengali (bn)
Gujarati (gu)
Hindi (hi)
Kannada (kn)
Marathi (mr)
Malayalam (ml)
Oriya (or)
Punjabi (pa)
Tamil (ta)
Telugu (te)
Task Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/mmlu-indic.gsm8k-indicolmOCR-Bench-English
olmOCR-bench (English Only)
This is a filtered version of the allenai/olmOCR-bench dataset containing only English documents.
Test Cases: Before vs After
Category
Before
After
Removed
% Retained
arxiv_math
2927
2917
10
99.7%
headers_footers
760
520
240
68.4%
long_tiny_text
442
442
0
100.0%
multi_column
884
691
193
78.2%
old_scans
526
526
0
100.0%
old_scans_math
458
458
0
100.0%
tables
1022
929
93
90.9%
TOTAL
7019
6483
536
92.4%
PDF Files:… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/olmOCR-Bench-English.samvaad-hi-v1100k high-quality conversations in English, Hindi, and Hinglish curated exclusively with an Indic context.
openswe-harbor
OpenSWE-Harbor — NeMo Gym ready
⚠️ Read this before training: there are TWO sets in here — full (45,316) and filtered (8,875).
Set
Tasks
Where it is
When to use
Full
45,316
routing/openswe_oss.jsonl, routing/openswe_other.jsonl, all of tasks/
Eval-only, dataset analysis, sweeps where you don't care about RL signal quality
Filtered (RL default)
8,875
routing/openswe_oss_filtered.jsonl, routing/openswe_other_filtered.jsonl, filtered_ids.txt
Use this for RL… See the full description on the dataset page: https://huggingface.co/datasets/ritvik-sarvam/openswe-harbor.
