CodeIsAbstract/sanskrit-samas-v1
Sanskrit Samas (Compound) Dataset — V1 Grammar-grounded training data for Sanskrit samas (compounds) covering all 7 samasa types, with laukik vigraha (natural paraphrase) and alaukik vigraha (Pāṇinian sUP analysis) on every row. Companion to sanskrit-sandhi-boundaries-v2 (sentence-level external sandhi). Merge both for a full sandhi+samas boundary training set. Dataset Rows: 258,408 Checksum: e7c0a804c109 Builder: benchmarks/build_samas_data.py (deterministic… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-samas-v1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face