CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BeitTigreAI /tigre-data-kenLM Tigre 5-gram Language Model (KenLM) Overview This repository provides a 5-gram Language Model (LM) for the Tigre language, trained using the KenLM toolkit. This model is a foundational resource for various downstream NLP and speech applications, including: Rescoring hypotheses in Automatic Speech Recognition (ASR). Improving text generation and fluency in Machine Translation (MT). Performing basic text filtering and quality control. The model is provided in the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-kenLM.0 likes47 downloads18d agoHugging Face02DaniilOr /humanized_correct_cleaning_kenlmtabular100K<n<1M0 likes41 downloads1y agoHugging Face03keystats /waxal-kenlm-models waxal-kenlm-models 5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark. These models are trained on normalized text — lowercased, with training targets restricted to a fixed alphabetic character set. This is the companion repo to… See the full description on the dataset page: https://huggingface.co/datasets/keystats/waxal-kenlm-models.0 likes38 downloads2mo agoHugging Face04MinhHan1009 /kenlmtext10M<n<100M0 likes37 downloads10mo agoHugging Face05pagantibet /Tibetan-KenLM-ACTibtrainingdata Tibetan Normalisation - KenLM ACTib Training Data A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.texttext-generation10M<n<100M0 likes37 downloads6mo agoHugging Face06keystats /waxal-kenlm-models-best waxal-kenlm-models-best 5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main-best CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark. These models are trained on non-normalized, raw text — case and punctuation are preserved, not lowercased or stripped. This is a deliberate choice… See the full description on the dataset page: https://huggingface.co/datasets/keystats/waxal-kenlm-models-best.0 likes37 downloads2mo agoHugging Face07Humphery7 /kenlm-datatext1K<n<10K0 likes32 downloads1y agoHugging Face08shahulyn /dhivehi-kenlm-corpus Dhivehi KenLM Language Model This repository contains a 5-gram KenLM language model built from the HuggingFace dataset: alakxender/dhivehi-news-corpus The corpus was cleaned and normalized to Thaana Unicode and used to improve Dhivehi ASR decoding (Wav2Vec2/XLS-R CTC). Files dhivehi.bin — KenLM binary language model (use this for decoding) dhivehi_corpus_raw.txt.gz — cleaned Dhivehi news text corpus Purpose This LM is designed to be used with CTC-based ASR… See the full description on the dataset page: https://huggingface.co/datasets/shahulyn/dhivehi-kenlm-corpus.text1M<n<10M0 likes17 downloads8mo agoHugging Face09chrisvinsen /id_kenlm_language_modeltext1K<n<10K0 likes10 downloads4y agoHugging Face10rupesh-pandey /hindi-kenlm-model0 likes9 downloads3mo agoHugging Face11marianna13 /the-eye-categories-kenlmtabular100K<n<1M0 likes5 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.