CoolFace
Datasetpublic

sanganaka/sanskrit_word_segmentation_dataset_2017

Sanskrit Word Segmentation and Morphological Candidate Dataset This dataset provides Sanskrit sentences annotated with gold-standard word segmentations, lemmas, and morphological tags based on the paper A Dataset for Sanskrit Word Segmentation. It also includes graph-based candidate morphological analyses derived from a structured Sanskrit parser. The dataset is useful for tasks such as: Word Segmentation Lemmatization Morphological Analysis Graph-based Disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/sanskrit_word_segmentation_dataset_2017.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes40downloads
filedata.jsonl623.1 MBdownload
fileDCS_pick.zip190.6 MBdownload
fileGraphml.zip620.6 MBdownload

sanganaka/sanskrit_word_segmentation_dataset_2017 · main · files are served by the source, never re-hosted here