CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tadad /midf-egangotri-sanskrit MIDF/eGangotri Sanskrit Manuscripts Reviewed line-segmentation annotations Segmentation v1.1 contains 2,879 reviewed pages with images, curved PAGE XML baselines, and editable geometry. Its 1,916 training pages contain 18,990 lines. A 60-page panel supports checkpoint selection, while 734 pages from three unseen manuscripts support broader validation. The test data contains 220 pages from the unseen M00638 manuscript and nine fixed adaptation pages from the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/midf-egangotri-sanskrit.documentimage-to-text100K<n<1M0 likes1.1k downloads9d agoHugging Face02addy88 /sanskrit-asr-84tabular10K<n<100K0 likes545 downloads5y agoHugging Face03NIVED47 /Sanskrit_ASR_Corpusaudio10K<n<100K1 likes411 downloads2y agoHugging Face04prathoshap /sushrota-sanskrit-asr-data Su-śrotā — Sanskrit ASR Dataset Curated and consented Sanskrit speech with utterance-level transcriptions, used to train the Su-śrotā Sanskrit ASR model (finetuned IndicConformer-CTC). Focused on śāstric and recitational Sanskrit (chant and prose). Author: Prof. Prathosh A P, Indian Institute of Science, Bengaluru. Audio: 16 kHz mono WAV. Transcriptions: Devanāgarī. Splits split clips hours description train 6,438 17.4 full training set (all sources… See the full description on the dataset page: https://huggingface.co/datasets/prathoshap/sushrota-sanskrit-asr-data.audioautomatic-speech-recognition1K<n<10K5 likes279 downloads27d agoHugging Face05surajp /shrutilipi_sanskrit Dataset Card for "shrutilipi_sanskrit" More Information needed audio10K<n<100K1 likes261 downloads3y agoHugging Face06addy88 /sanskrit-asr-84-evaltabular1K<n<10K1 likes227 downloads5y agoHugging Face07SanskritVoyager /SanskritTravelogue Sanskrit Travelogue: Unified Sanskrit Text Corpus Unified, deduplicated, and morphologically annotated Sanskrit corpus, aggregating 13,010 texts (~182 million words, ~15.7 million segments) from 8 major digital Sanskrit libraries. All texts are normalized to IAST (International Alphabet of Sanskrit Transliteration). Note: Some annotations are still missing; In a follow up release, there will be added paragraph level metadata and translations for the GRETIL and SARIT corpus, plus… See the full description on the dataset page: https://huggingface.co/datasets/SanskritVoyager/SanskritTravelogue.tabulartext-classification10M<n<100M0 likes210 downloads3mo agoHugging Face08meharuhanzz /OCR-Bench1000-Sanskrit OCR-Bench1000-Sanskrit 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Sanskrit OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category sanskrit_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Sanskrit.imageimage-to-text1K<n<10K0 likes152 downloads9d agoHugging Face09pavanmantha /sanskrit_asraudio10K<n<100K0 likes120 downloads7mo agoHugging Face10bpHigh /iNLTK_Sanskrit_Shlokas_Datasettextn<1K2 likes105 downloads2y agoHugging Face11pari-kulkarni /itihasa-sanskrit-en-filtered Itihasa Sanskrit-English Parallel Corpus (Quality-Filtered) A cleaned, deduplicated, leakage-audited version of the Itihasa corpus (Aralikatte et al., 2021), containing Sanskrit-English verse pairs from the Ramayana and Mahabharata. Prepared as Phase 1 of the SALIDLab Sanskrit project. Corpus statistics Split Pairs Sanskrit tokens English tokens Avg SA len Avg EN len train 74,650 833,331 2,294,746 11.16 30.74 dev 6,137 70,048 192,430 11.41 31.36… See the full description on the dataset page: https://huggingface.co/datasets/pari-kulkarni/itihasa-sanskrit-en-filtered.texttranslation10K<n<100K1 likes87 downloads3mo agoHugging Face12CodeIsAbstract /sanskrit-sandhi-boundaries-v2 Sanskrit Sandhi Boundary Dataset (V3 — verified, category-complete) Training data for the sandhi boundary-detection model in CodeIsAbstract/sanskrit-sandhi-boundary-v2. The task: given a sandhi-joined string (a compound or multi-word string), predict the character positions where independent words end, so a downstream Sanskrit tokenizer can split it into complete, independent tokens. This is the verified release: every row has been passed through a deterministic sanitizer… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-boundaries-v2.texttoken-classification1M<n<10M0 likes86 downloads29d agoHugging Face13Anamavajra-Labs /sanskrit-karaka-hypergraph Sanskrit Kāraka Hypergraph A predication hypergraph over Sanskrit: vertices are lemma types, hyperedges are predications, and each tine carries a Pāṇinian kāraka role. Why a hypergraph rather than a graph of binary relations: a sentence is an n-ary predicate, and an n-ary relation does not survive projection onto its binary sub-relations. Given only the pairs agent–object, object–recipient and agent–recipient you can no longer tell whether there was one three-place act or three… See the full description on the dataset page: https://huggingface.co/datasets/Anamavajra-Labs/sanskrit-karaka-hypergraph.tabulartoken-classification100K<n<1M0 likes82 downloads2mo agoHugging Face14paws /sanskrit-verses-gretil Sanskrit Literature Source Retrieval Dataset (GRETIL) This dataset contains 283,935 Sanskrit verses from the GRETIL (Göttingen Register of Electronic Texts in Indian Languages) corpus, designed for training language models on Sanskrit literature source identification. Dataset Description This dataset is created for a reinforcement learning task where models learn to identify the source of Sanskrit quotes, including: Genre (kavya, epic, purana, veda, shastra, tantra… See the full description on the dataset page: https://huggingface.co/datasets/paws/sanskrit-verses-gretil.texttext-generation100K<n<1M0 likes81 downloads1y agoHugging Face15chronbmm /sanskrit-multitask-devanagaritext1M<n<10M0 likes77 downloads2y agoHugging Face16harsha-desaraju /telugu-sanskrit-english-texttext1M<n<10M0 likes77 downloads3mo agoHugging Face17chronbmm /sanskrit-sandhi-split-sighum Dataset Card for "sanskrit-sandhi-split-sighum" More Information needed text100K<n<1M0 likes69 downloads3y agoHugging Face18chronbmm /sanskrit-multitasktext1M<n<10M0 likes68 downloads2y agoHugging Face19saillab /testalpaca_sanskrit_tacoThis repo consists of the datasets used for the TaCo paper. There are four datasets: Multilingual Alpaca-52K GPT-4 dataset Multilingual Dolly-15K GPT-4 dataset TaCo dataset Multilingual Vicuna Benchmark dataset We translated the first three datasets using Google Cloud Translation. The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets. If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/testalpaca_sanskrit_taco.text10K<n<100K0 likes67 downloads2y agoHugging Face20vjdevane /sanskriti Citation @misc{maji2025sanskriticomprehensivebenchmarkevaluating, title={SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture}, author={Arijit Maji and Raghvendra Kumar and Akash Ghosh and Anushka and Sriparna Saha}, year={2025}, eprint={2506.15355}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2506.15355}, } textmultiple-choice10K<n<100K0 likes66 downloads1y agoHugging Face21cfilt /RoundTripOCR-sanskritPost-OCR error correction dataset (train, test and validation set) for Sanskrit language generated using RoundTripOCR technique. Code: https://github.com/harshvivek14/RoundTripOCR text1M<n<10M1 likes65 downloads2y agoHugging Face22shunyasea /vedic-sanskrit Dataset Card for "vedic-sanskrit" More Information needed text100K<n<1M0 likes60 downloads3y agoHugging Face23CodeIsAbstract /sanskrit-morpho-sequences Sanskrit Morphological Sequence Corpus (Vidyut-Verified) A large-scale, Pāṇinian-verified morphological sequence dataset for classical and Vedic Sanskrit. Every token is annotated with its lemma, generative root (aupadeśika), part-of-speech, case, number, person, voice, and gender — all in the SLP1 transliteration, and all aligned at the sentence level for sequence-tagging / seq2seq training. 710,785 sentences (after deduplication) 5,511,664 tokens 14 columns (10 linguistic + 4… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-morpho-sequences.tabulartoken-classification100K<n<1M0 likes59 downloads2mo agoHugging Face24harsha-desaraju /sanskrit-text-telugu-scripttext100K<n<1M0 likes57 downloads3mo agoHugging Face25Process-Venue /Sanskrit-OCR-Typed-Dataset Sanskrit OCR Dataset This dataset contains Sanskrit text images paired with their corresponding text labels, designed for OCR (Optical Character Recognition) tasks. Dataset Structure The dataset is split into training and validation sets: Training set: Contains unique Sanskrit text images Validation set: Contains separate unique Sanskrit text images Features image: The image containing Sanskrit text label: The corresponding Sanskrit text label filename:… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-OCR-Typed-Dataset.imageimage-classification1K<n<10K2 likes56 downloads2y agoHugging Face26PleIAs /Sanskrit-PDtabular1K<n<10K2 likes55 downloads2y agoHugging Face27chronbmm /sanskrit-unsandhi-morphosyntax-taggingtext100K<n<1M0 likes51 downloads2y agoHugging Face2813ari /Sanskriti SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture Link: arxiv.org/abs/2506.15355 Dataset Description The SANSKRITI benchmark is the largest dataset created to evaluate Language Models' (LMs) comprehension and reasoning capabilities regarding the rich cultural diversity of India.It addresses the critical need for culturally-aware benchmarks, as the global effectiveness of LMs depends on their understanding of local… See the full description on the dataset page: https://huggingface.co/datasets/13ari/Sanskriti.textquestion-answering10K<n<100K2 likes50 downloads11mo agoHugging Face29chronbmm /sanskrit-monolingual-pretraining Dataset Card for "sanskrit-monolingual-pretraining" More Information needed text10M<n<100M2 likes47 downloads3y agoHugging Face30CodeIsAbstract /sanskrit-samas-v1 Sanskrit Samas (Compound) Dataset — V1 Grammar-grounded training data for Sanskrit samas (compounds) covering all 7 samasa types, with laukik vigraha (natural paraphrase) and alaukik vigraha (Pāṇinian sUP analysis) on every row. Companion to sanskrit-sandhi-boundaries-v2 (sentence-level external sandhi). Merge both for a full sandhi+samas boundary training set. Dataset Rows: 258,408 Checksum: e7c0a804c109 Builder: benchmarks/build_samas_data.py (deterministic… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-samas-v1.texttoken-classification100K<n<1M0 likes47 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.