datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tigre-data-kenLM
Tigre 5-gram Language Model (KenLM)
Overview
This repository provides a 5-gram Language Model (LM) for the Tigre language, trained using the KenLM toolkit. This model is a foundational resource for various downstream NLP and speech applications, including:
Rescoring hypotheses in Automatic Speech Recognition (ASR).
Improving text generation and fluency in Machine Translation (MT).
Performing basic text filtering and quality control.
The model is provided in the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-kenLM.humanized_correct_cleaning_kenlmwaxal-kenlm-models
waxal-kenlm-models
5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark.
These models are trained on normalized text — lowercased, with training targets restricted to a fixed alphabetic character set. This is the companion repo to… See the full description on the dataset page: https://huggingface.co/datasets/keystats/waxal-kenlm-models.kenlmTibetan-KenLM-ACTibtrainingdata
Tibetan Normalisation - KenLM ACTib Training Data
A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.waxal-kenlm-models-best
waxal-kenlm-models-best
5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main-best CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark.
These models are trained on non-normalized, raw text — case and punctuation are preserved, not lowercased or stripped. This is a deliberate choice… See the full description on the dataset page: https://huggingface.co/datasets/keystats/waxal-kenlm-models-best.kenlm-datadhivehi-kenlm-corpus
Dhivehi KenLM Language Model
This repository contains a 5-gram KenLM language model built from the HuggingFace dataset:
alakxender/dhivehi-news-corpus
The corpus was cleaned and normalized to Thaana Unicode and used to improve Dhivehi ASR decoding (Wav2Vec2/XLS-R CTC).
Files
dhivehi.bin — KenLM binary language model (use this for decoding)
dhivehi_corpus_raw.txt.gz — cleaned Dhivehi news text corpus
Purpose
This LM is designed to be used with CTC-based ASR… See the full description on the dataset page: https://huggingface.co/datasets/shahulyn/dhivehi-kenlm-corpus.id_kenlm_language_modelhindi-kenlm-modelthe-eye-categories-kenlm
