CoolFace
Datasetpublic

shahulyn/dhivehi-kenlm-corpus

Dhivehi KenLM Language Model This repository contains a 5-gram KenLM language model built from the HuggingFace dataset: alakxender/dhivehi-news-corpus The corpus was cleaned and normalized to Thaana Unicode and used to improve Dhivehi ASR decoding (Wav2Vec2/XLS-R CTC). Files dhivehi.bin — KenLM binary language model (use this for decoding) dhivehi_corpus_raw.txt.gz — cleaned Dhivehi news text corpus Purpose This LM is designed to be used with… See the full description on the dataset page: https://huggingface.co/datasets/shahulyn/dhivehi-kenlm-corpus.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes16downloads
filedhivehi_corpus_raw.txt.gz51.6 MBdownload
filedhivehi.bin209.1 MBdownload

shahulyn/dhivehi-kenlm-corpus · main · files are served by the source, never re-hosted here