CoolFace
20 results

sandhi

CodeIsAbstract /sanskrit-sandhi-boundaries-v2 Sanskrit Sandhi Boundary Dataset (V3 — verified, category-complete) Training data for the sandhi boundary-detection model in CodeIsAbstract/sanskrit-sandhi-boundary-v2. The task: given a sandhi-joined string (a compound or multi-word string), predict the character positions where independent words end, so a downstream Sanskrit tokenizer can split it into complete, independent tokens. This is the verified release: every row has been passed through a deterministic sanitizer… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-boundaries-v2.texttoken-classification1M<n<10M0 likes86 downloads1mo agoHugging Facechronbmm /sanskrit-sandhi-split-sighum Dataset Card for "sanskrit-sandhi-split-sighum" More Information needed text100K<n<1M0 likes69 downloads3y agoHugging Facechronbmm /sanskrit-sandhi-split-hackathon Dataset Card for "sanskrit-sandhi-split-hackathon" More Information needed text100K<n<1M0 likes44 downloads3y agoHugging Facenithinmanoj10 /Sandhi-Morph-1.0 Sandhi-Morph-1.0 Malayalam word analyses: 921,531 training rows, plus five held-out evaluation sets. It accompanies the model Sandhi-1.0-8M. Each row is one word (not a sentence) with its sandhi split, morphemes and their grammatical function, root lemma, grammatical features, and for some rows IPA and syllables. Configs and splits Config Split Rows What default train 903,078 Training rows default dev 18,453 Development rows. No word appears in both… See the full description on the dataset page: https://huggingface.co/datasets/nithinmanoj10/Sandhi-Morph-1.0.tabulartoken-classification100K<n<1M0 likes44 downloads5d agoHugging FaceCodeIsAbstract /sanskrit-sandhi-samas-v3 Sanskrit Sandhi + Samas Boundary Dataset — V3 The canonical merged training dataset for the sanskrit-sandhi-boundary model: sentence-level external sandhi (corpus) + grammar-generated samas (compounds, all 7 types) in one file. Composition Source Rows Description sandhi_corpus 741,803 running-text word boundaries (external sandhi) from sanskrit-sandhi-boundaries-v2 samas 258,408 grammar-generated compounds (7 types, laukik + alaukik vigraha) from… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-samas-v3.token-classification0 likes40 downloads28d agoHugging Facechronbmm /sandhi-split-2018text100K<n<1M0 likes29 downloads2y agoHugging Face