CoolFace
17 results

kenlm

BeitTigreAI /tigre-data-kenLM Tigre 5-gram Language Model (KenLM) Overview This repository provides a 5-gram Language Model (LM) for the Tigre language, trained using the KenLM toolkit. This model is a foundational resource for various downstream NLP and speech applications, including: Rescoring hypotheses in Automatic Speech Recognition (ASR). Improving text generation and fluency in Machine Translation (MT). Performing basic text filtering and quality control. The model is provided in the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-kenLM.0 likes47 downloads18d agoHugging FaceDaniilOr /humanized_correct_cleaning_kenlmtabular100K<n<1M0 likes41 downloads1y agoHugging Facekeystats /waxal-kenlm-models waxal-kenlm-models 5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark. These models are trained on normalized text — lowercased, with training targets restricted to a fixed alphabetic character set. This is the companion repo to… See the full description on the dataset page: https://huggingface.co/datasets/keystats/waxal-kenlm-models.0 likes38 downloads2mo agoHugging FaceMinhHan1009 /kenlmtext10M<n<100M0 likes37 downloads10mo agoHugging Facepagantibet /Tibetan-KenLM-ACTibtrainingdata Tibetan Normalisation - KenLM ACTib Training Data A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.texttext-generation10M<n<100M0 likes37 downloads6mo agoHugging Facekeystats /waxal-kenlm-models-best waxal-kenlm-models-best 5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main-best CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark. These models are trained on non-normalized, raw text — case and punctuation are preserved, not lowercased or stripped. This is a deliberate choice… See the full description on the dataset page: https://huggingface.co/datasets/keystats/waxal-kenlm-models-best.0 likes37 downloads2mo agoHugging Face