CoolFace
13 results

word-segmentation

KhoaUIT /VNDS-Vietnews-WordSegmentationtext100K<n<1M0 likes34 downloads2y agoHugging Facebillingsmoore /tibetan-word-segmentation-ds Tibetan Word Segmentation Annotations Human expert word boundary annotations for 100 utterances of modern Lhasa Tibetan news speech, drawn from the NICT-Tib1 ASR corpus. Dataset Description Tibetan orthography does not mark spaces between words; syllables are delimited only by the tsek (་) character. This dataset provides human-annotated word segmentations that serve as the reference standard in a five-way segmentation agreement study comparing automatic Tibetan… See the full description on the dataset page: https://huggingface.co/datasets/billingsmoore/tibetan-word-segmentation-ds.texttoken-classificationn<1K0 likes30 downloads2mo agoHugging Facesanganaka /sanskrit_word_segmentation_dataset_2017 Sanskrit Word Segmentation and Morphological Candidate Dataset This dataset provides Sanskrit sentences annotated with gold-standard word segmentations, lemmas, and morphological tags based on the paper A Dataset for Sanskrit Word Segmentation. It also includes graph-based candidate morphological analyses derived from a structured Sanskrit parser. The dataset is useful for tasks such as: Word Segmentation Lemmatization Morphological Analysis Graph-based Disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/sanskrit_word_segmentation_dataset_2017.text100K<n<1M0 likes28 downloads1y agoHugging Facephonsobon /khmer-word-segmentationgated Khmer Administrative Text Dataset 🇰🇭 Overview This dataset contains Khmer-language sentences that reflect formal administrative and government-style writing. The dataset was synthetically generated using the Gemini large language model, developed by Google, to simulate official Khmer document language such as reports, letters, and institutional communication. Data Source Transparency Source: Synthetic data generated by Gemini (Google) Type:… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/khmer-word-segmentation.text100K<n<1M0 likes3 downloads5mo agoHugging Face