datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humanized_correct_cleaning_kenlmkenlmTibetan-KenLM-ACTibtrainingdata
Tibetan Normalisation - KenLM ACTib Training Data
A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.kenlm-datadhivehi-kenlm-corpus
Dhivehi KenLM Language Model
This repository contains a 5-gram KenLM language model built from the HuggingFace dataset:
alakxender/dhivehi-news-corpus
The corpus was cleaned and normalized to Thaana Unicode and used to improve Dhivehi ASR decoding (Wav2Vec2/XLS-R CTC).
Files
dhivehi.bin — KenLM binary language model (use this for decoding)
dhivehi_corpus_raw.txt.gz — cleaned Dhivehi news text corpus
Purpose
This LM is designed to be used with CTC-based ASR… See the full description on the dataset page: https://huggingface.co/datasets/shahulyn/dhivehi-kenlm-corpus.id_kenlm_language_modelthe-eye-categories-kenlm
