CoolFace
Datasetpublic

CodeIsAbstract/sanskrit-samas-v1

Sanskrit Samas (Compound) Dataset — V1 Grammar-grounded training data for Sanskrit samas (compounds) covering all 7 samasa types, with laukik vigraha (natural paraphrase) and alaukik vigraha (Pāṇinian sUP analysis) on every row. Companion to sanskrit-sandhi-boundaries-v2 (sentence-level external sandhi). Merge both for a full sandhi+samas boundary training set. Dataset Rows: 258,408 Checksum: e7c0a804c109 Builder: benchmarks/build_samas_data.py (deterministic… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-samas-v1.

sourceHugging Facemitupdated 29d agoView on Hugging Face
0likes47downloads
3 commits on main
e1c211a29d ago

Upload data/samas_V1.jsonl with huggingface_hub

CodeIsAbstract
62908bc29d ago

Upload README.md with huggingface_hub

CodeIsAbstract
9b0ba7529d ago

initial commit

CodeIsAbstract