datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
african-transcribed-speech-monolingual
African Transcribed Speech — Monolingual
Sentence-level monolingual text for 15 African languages, derived from translated religious speech transcriptions. Intended as reference text for evaluating speech machine translation (e.g. BLEU scoring), and as a monolingual corpus for language modeling / tokenizer training.
Each language is a separate subset — load with e.g. load_dataset("<repo>", "fat").
Coverage
Language
Code
Sentences
Malagasy
mlg
108,333… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/african-transcribed-speech-monolingual.manipuri-monolingual-corpus
Manipuri Monolingual Corpus
About
This Manipuri Monolingual corpus contains an expanded monolingual corpus for Manipuri in the following paper. It has been compiled from publicly available texts on the internet in the open domain.
Dataset Statistics
Set 1 Dataset contains approx. 11 million words.
Set 2 Dataset contains approx. 19 million words.
Set 3 Dataset contains approx. 76 million words.
Dataset Quality
Set 1 is of high quality. Set 3 is of low… See the full description on the dataset page: https://huggingface.co/datasets/joyson117/manipuri-monolingual-corpus.
