KenLM
Datasets
All datasets matching “KenLM”Tibetan-KenLM-ACTibtrainingdata
Tibetan Normalisation - KenLM ACTib Training Data
A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.tigre-data-kenLM
Tigre 5-gram Language Model (KenLM)
Overview
This repository provides a 5-gram Language Model (LM) for the Tigre language, trained using the KenLM toolkit. This model is a foundational resource for various downstream NLP and speech applications, including:
Rescoring hypotheses in Automatic Speech Recognition (ASR).
Improving text generation and fluency in Machine Translation (MT).
Performing basic text filtering and quality control.
The model is provided in the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-kenLM.waxal-kenlm-models
waxal-kenlm-models
5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark.
These models are trained on normalized text — lowercased, with training targets restricted to a fixed alphabetic character set. This is the companion repo to… See the full description on the dataset page: https://huggingface.co/datasets/keystats/waxal-kenlm-models.humanized_correct_cleaning_kenlmkenlmkenlm-data
