word segmentation
VNDS-Vietnews-WordSegmentationtibetan-word-segmentation-ds
Tibetan Word Segmentation Annotations
Human expert word boundary annotations for 100 utterances of modern Lhasa Tibetan news speech, drawn from the NICT-Tib1 ASR corpus.
Dataset Description
Tibetan orthography does not mark spaces between words; syllables are delimited only by the tsek (་) character. This dataset provides human-annotated word segmentations that serve as the reference standard in a five-way segmentation agreement study comparing automatic Tibetan… See the full description on the dataset page: https://huggingface.co/datasets/billingsmoore/tibetan-word-segmentation-ds.sanskrit_word_segmentation_dataset_2017
Sanskrit Word Segmentation and Morphological Candidate Dataset
This dataset provides Sanskrit sentences annotated with gold-standard word segmentations, lemmas, and morphological tags based on the paper A Dataset for Sanskrit Word Segmentation. It also includes graph-based candidate morphological analyses derived from a structured Sanskrit parser. The dataset is useful for tasks such as:
Word Segmentation
Lemmatization
Morphological Analysis
Graph-based Disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/sanskrit_word_segmentation_dataset_2017.khmer-word-segmentation
Khmer Administrative Text Dataset 🇰🇭
Overview
This dataset contains Khmer-language sentences that reflect formal administrative and government-style writing.
The dataset was synthetically generated using the Gemini large language model, developed by Google, to simulate official Khmer document language such as reports, letters, and institutional communication.
Data Source Transparency
Source: Synthetic data generated by Gemini (Google)
Type:… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/khmer-word-segmentation.
