sanganaka/sanskrit_word_segmentation_dataset_2017
Sanskrit Word Segmentation and Morphological Candidate Dataset This dataset provides Sanskrit sentences annotated with gold-standard word segmentations, lemmas, and morphological tags based on the paper A Dataset for Sanskrit Word Segmentation. It also includes graph-based candidate morphological analyses derived from a structured Sanskrit parser. The dataset is useful for tasks such as: Word Segmentation Lemmatization Morphological Analysis Graph-based Disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/sanskrit_word_segmentation_dataset_2017.
040
