sanskrit
Datasets
All datasets matching “sanskrit”sanskrit_books_collectionmidf-egangotri-sanskrit
MIDF/eGangotri Sanskrit Manuscripts
Reviewed line-segmentation annotations
Segmentation v1.1 contains 2,879 reviewed
pages with images, curved PAGE XML baselines, and editable geometry. Its 1,916
training pages contain 18,990 lines. A 60-page panel supports checkpoint
selection, while 734 pages from three unseen manuscripts support broader
validation. The test data contains 220 pages from the unseen M00638 manuscript
and nine fixed adaptation pages from the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/midf-egangotri-sanskrit.sanskrit-asr-84Sanskrit_ASR_Corpussample_sanskrit_pile
Sanskrit Pile v0
A tokenizer-agnostic, streamable raw pretraining corpus of Devanagari Sanskrit text,
assembled for continual pretraining of Sanskrit LLMs (1B–3B proof-of-concept).
Stats
Documents: 1,429,519
Characters: ~6.19 Billion
Approx tokens (~4 chars/tok): ~1.55 Billion
Format: sharded Parquet (zstd), FineWeb-style
Script: Devanagari (IAST/SLP1/ITRANS normalized via indic-transliteration)
Schema (tokenizer-agnostic)
field
type… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sample_sanskrit_pile.sanskrit-ocr-post-correction\
A Benchmark and Dataset for Post-OCR text correction in Sanskrit.
This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script.
It includes:
- Train/Validation/Test splits with OCR text and corrected ground truth
- An out-of-domain test set of 500 sentences
- Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and Goladhyaya
