CoolFace
Datasetpublic

CodeIsAbstract/sanskrit-sandhi-samas-v3

Sanskrit Sandhi + Samas Boundary Dataset — V3 The canonical merged training dataset for the sanskrit-sandhi-boundary model: sentence-level external sandhi (corpus) + grammar-generated samas (compounds, all 7 types) in one file. Composition Source Rows Description sandhi_corpus 741,803 running-text word boundaries (external sandhi) from sanskrit-sandhi-boundaries-v2 samas 258,408 grammar-generated compounds (7 types, laukik + alaukik vigraha) from… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-samas-v3.

sourceHugging Facemitupdated 29d agoView on Hugging Face
0likes40downloads
settings

This repository belongs to CodeIsAbstract on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesanskrit-sandhi-samas-v3
visibilitypublic
licencemit
gatedno
ownerCodeIsAbstract
Account settings
CodeIsAbstract/sanskrit-sandhi-samas-v3 · CoolFace