CoolFace
Datasetpublic

CodeIsAbstract/sanskrit-sandhi-samas-v3

Sanskrit Sandhi + Samas Boundary Dataset — V3 The canonical merged training dataset for the sanskrit-sandhi-boundary model: sentence-level external sandhi (corpus) + grammar-generated samas (compounds, all 7 types) in one file. Composition Source Rows Description sandhi_corpus 741,803 running-text word boundaries (external sandhi) from sanskrit-sandhi-boundaries-v2 samas 258,408 grammar-generated compounds (7 types, laukik + alaukik vigraha) from… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-samas-v3.

sourceHugging Facemitupdated 29d agoView on Hugging Face
0likes40downloads
3 commits on main
94a03b229d ago

Upload data/train_V3_sandhi_samas.jsonl with huggingface_hub

CodeIsAbstract
71aaf0329d ago

Upload README.md with huggingface_hub

CodeIsAbstract
f4fa18629d ago

initial commit

CodeIsAbstract