CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DaniilOr /humanized_correct_cleaning_kenlmtabular100K<n<1M0 likes41 downloads1y agoHugging Face02MinhHan1009 /kenlmtext10M<n<100M0 likes37 downloads10mo agoHugging Face03pagantibet /Tibetan-KenLM-ACTibtrainingdata Tibetan Normalisation - KenLM ACTib Training Data A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.texttext-generation10M<n<100M0 likes37 downloads6mo agoHugging Face04Humphery7 /kenlm-datatext1K<n<10K0 likes32 downloads1y agoHugging Face05shahulyn /dhivehi-kenlm-corpus Dhivehi KenLM Language Model This repository contains a 5-gram KenLM language model built from the HuggingFace dataset: alakxender/dhivehi-news-corpus The corpus was cleaned and normalized to Thaana Unicode and used to improve Dhivehi ASR decoding (Wav2Vec2/XLS-R CTC). Files dhivehi.bin — KenLM binary language model (use this for decoding) dhivehi_corpus_raw.txt.gz — cleaned Dhivehi news text corpus Purpose This LM is designed to be used with CTC-based ASR… See the full description on the dataset page: https://huggingface.co/datasets/shahulyn/dhivehi-kenlm-corpus.text1M<n<10M0 likes17 downloads8mo agoHugging Face06chrisvinsen /id_kenlm_language_modeltext1K<n<10K0 likes10 downloads4y agoHugging Face07marianna13 /the-eye-categories-kenlmtabular100K<n<1M0 likes5 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.