CoolFace
Datasetpublic

CodeIsAbstract/sanskrit-samas-v1

Sanskrit Samas (Compound) Dataset — V1 Grammar-grounded training data for Sanskrit samas (compounds) covering all 7 samasa types, with laukik vigraha (natural paraphrase) and alaukik vigraha (Pāṇinian sUP analysis) on every row. Companion to sanskrit-sandhi-boundaries-v2 (sentence-level external sandhi). Merge both for a full sandhi+samas boundary training set. Dataset Rows: 258,408 Checksum: e7c0a804c109 Builder: benchmarks/build_samas_data.py (deterministic… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-samas-v1.

sourceHugging Facemitupdated 29d agoView on Hugging Face
0likes47downloads
settings

This repository belongs to CodeIsAbstract on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesanskrit-samas-v1
visibilitypublic
licencemit
gatedno
ownerCodeIsAbstract
Account settings
CodeIsAbstract/sanskrit-samas-v1 · CoolFace