datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midf-egangotri-sanskrit
MIDF/eGangotri Sanskrit Manuscripts
Reviewed line-segmentation annotations
Segmentation v1.1 contains 2,879 reviewed
pages with images, curved PAGE XML baselines, and editable geometry. Its 1,916
training pages contain 18,990 lines. A 60-page panel supports checkpoint
selection, while 734 pages from three unseen manuscripts support broader
validation. The test data contains 220 pages from the unseen M00638 manuscript
and nine fixed adaptation pages from the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/midf-egangotri-sanskrit.sanskrit-asr-84Sanskrit_ASR_Corpussushrota-sanskrit-asr-data
Su-śrotā — Sanskrit ASR Dataset
Curated and consented Sanskrit speech with utterance-level transcriptions, used to train the
Su-śrotā Sanskrit ASR model
(finetuned IndicConformer-CTC). Focused on śāstric and recitational Sanskrit (chant and prose).
Author: Prof. Prathosh A P, Indian Institute of Science, Bengaluru.
Audio: 16 kHz mono WAV. Transcriptions: Devanāgarī.
Splits
split
clips
hours
description
train
6,438
17.4
full training set (all sources… See the full description on the dataset page: https://huggingface.co/datasets/prathoshap/sushrota-sanskrit-asr-data.shrutilipi_sanskrit
Dataset Card for "shrutilipi_sanskrit"
More Information needed
sanskrit-asr-84-evalSanskritTravelogue
Sanskrit Travelogue: Unified Sanskrit Text Corpus
Unified, deduplicated, and morphologically annotated Sanskrit corpus, aggregating 13,010 texts (~182 million words, ~15.7 million segments) from 8 major digital Sanskrit libraries. All texts are normalized to IAST (International Alphabet of Sanskrit Transliteration).
Note:
Some annotations are still missing;
In a follow up release, there will be added paragraph level metadata and translations for the GRETIL and SARIT corpus, plus… See the full description on the dataset page: https://huggingface.co/datasets/SanskritVoyager/SanskritTravelogue.OCR-Bench1000-Sanskrit
OCR-Bench1000-Sanskrit
1000 synthetic printed-text line images with ground-truth transcriptions,
sampled from a larger locally-held Sanskrit OCR training corpus.
This is a benchmark/sample release, not the full training set.
Data fields
Field
Description
file_name
relative path to the image (images/...)
text
ground-truth transcription
category
sanskrit_only / english_only / mixed / numeric_and_symbols
length_bucket
short / medium / long, by… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Sanskrit.sanskrit_asriNLTK_Sanskrit_Shlokas_Datasetitihasa-sanskrit-en-filtered
Itihasa Sanskrit-English Parallel Corpus (Quality-Filtered)
A cleaned, deduplicated, leakage-audited version of the Itihasa corpus (Aralikatte et al., 2021), containing Sanskrit-English verse pairs from the Ramayana and Mahabharata.
Prepared as Phase 1 of the SALIDLab Sanskrit project.
Corpus statistics
Split
Pairs
Sanskrit tokens
English tokens
Avg SA len
Avg EN len
train
74,650
833,331
2,294,746
11.16
30.74
dev
6,137
70,048
192,430
11.41
31.36… See the full description on the dataset page: https://huggingface.co/datasets/pari-kulkarni/itihasa-sanskrit-en-filtered.sanskrit-sandhi-boundaries-v2
Sanskrit Sandhi Boundary Dataset (V3 — verified, category-complete)
Training data for the sandhi boundary-detection model in
CodeIsAbstract/sanskrit-sandhi-boundary-v2.
The task: given a sandhi-joined string (a compound or multi-word string),
predict the character positions where independent words end, so a downstream
Sanskrit tokenizer can split it into complete, independent tokens.
This is the verified release: every row has been passed through a
deterministic sanitizer… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-boundaries-v2.sanskrit-karaka-hypergraph
Sanskrit Kāraka Hypergraph
A predication hypergraph over Sanskrit: vertices are lemma types, hyperedges are
predications, and each tine carries a Pāṇinian kāraka role.
Why a hypergraph rather than a graph of binary relations: a sentence is an n-ary
predicate, and an n-ary relation does not survive projection onto its binary
sub-relations. Given only the pairs agent–object, object–recipient and
agent–recipient you can no longer tell whether there was one three-place act or
three… See the full description on the dataset page: https://huggingface.co/datasets/Anamavajra-Labs/sanskrit-karaka-hypergraph.sanskrit-verses-gretil
Sanskrit Literature Source Retrieval Dataset (GRETIL)
This dataset contains 283,935 Sanskrit verses from the GRETIL (Göttingen Register of Electronic Texts in Indian Languages) corpus, designed for training language models on Sanskrit literature source identification.
Dataset Description
This dataset is created for a reinforcement learning task where models learn to identify the source of Sanskrit quotes, including:
Genre (kavya, epic, purana, veda, shastra, tantra… See the full description on the dataset page: https://huggingface.co/datasets/paws/sanskrit-verses-gretil.sanskrit-multitask-devanagaritelugu-sanskrit-english-textsanskrit-sandhi-split-sighum
Dataset Card for "sanskrit-sandhi-split-sighum"
More Information needed
sanskrit-multitasktestalpaca_sanskrit_tacoThis repo consists of the datasets used for the TaCo paper. There are four datasets:
Multilingual Alpaca-52K GPT-4 dataset
Multilingual Dolly-15K GPT-4 dataset
TaCo dataset
Multilingual Vicuna Benchmark dataset
We translated the first three datasets using Google Cloud Translation.
The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets.
If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/testalpaca_sanskrit_taco.sanskriti
Citation
@misc{maji2025sanskriticomprehensivebenchmarkevaluating,
title={SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture},
author={Arijit Maji and Raghvendra Kumar and Akash Ghosh and Anushka and Sriparna Saha},
year={2025},
eprint={2506.15355},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.15355},
}
RoundTripOCR-sanskritPost-OCR error correction dataset (train, test and validation set) for Sanskrit language generated using RoundTripOCR technique.
Code: https://github.com/harshvivek14/RoundTripOCR
vedic-sanskrit
Dataset Card for "vedic-sanskrit"
More Information needed
sanskrit-morpho-sequences
Sanskrit Morphological Sequence Corpus (Vidyut-Verified)
A large-scale, Pāṇinian-verified morphological sequence dataset for
classical and Vedic Sanskrit. Every token is annotated with its lemma,
generative root (aupadeśika), part-of-speech, case, number, person, voice,
and gender — all in the SLP1 transliteration, and all aligned at the
sentence level for sequence-tagging / seq2seq training.
710,785 sentences (after deduplication)
5,511,664 tokens
14 columns (10 linguistic + 4… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-morpho-sequences.sanskrit-text-telugu-scriptSanskrit-OCR-Typed-Dataset
Sanskrit OCR Dataset
This dataset contains Sanskrit text images paired with their corresponding text labels, designed for OCR (Optical Character Recognition) tasks.
Dataset Structure
The dataset is split into training and validation sets:
Training set: Contains unique Sanskrit text images
Validation set: Contains separate unique Sanskrit text images
Features
image: The image containing Sanskrit text
label: The corresponding Sanskrit text label
filename:… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-OCR-Typed-Dataset.Sanskrit-PDsanskrit-unsandhi-morphosyntax-taggingSanskriti
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
Link: arxiv.org/abs/2506.15355
Dataset Description
The SANSKRITI benchmark is the largest dataset created to evaluate Language Models' (LMs) comprehension and reasoning capabilities regarding the rich cultural diversity of India.It addresses the critical need for culturally-aware benchmarks, as the global effectiveness of LMs depends on their understanding of local… See the full description on the dataset page: https://huggingface.co/datasets/13ari/Sanskriti.sanskrit-monolingual-pretraining
Dataset Card for "sanskrit-monolingual-pretraining"
More Information needed
sanskrit-samas-v1
Sanskrit Samas (Compound) Dataset — V1
Grammar-grounded training data for Sanskrit samas (compounds) covering all
7 samasa types, with laukik vigraha (natural paraphrase) and alaukik
vigraha (Pāṇinian sUP analysis) on every row.
Companion to sanskrit-sandhi-boundaries-v2
(sentence-level external sandhi). Merge both for a full sandhi+samas boundary
training set.
Dataset
Rows: 258,408
Checksum: e7c0a804c109
Builder: benchmarks/build_samas_data.py (deterministic… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-samas-v1.
