datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
south-african-monolingual-corpora-jsonl
South African Languages Pretraining Dataset
This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections.
The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity
Languages Included
Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.telugu-monolingual-datasetafrican-transcribed-speech-monolingual
African Transcribed Speech — Monolingual
Sentence-level monolingual text for 15 African languages, derived from translated religious speech transcriptions. Intended as reference text for evaluating speech machine translation (e.g. BLEU scoring), and as a monolingual corpus for language modeling / tokenizer training.
Each language is a separate subset — load with e.g. load_dataset("<repo>", "fat").
Coverage
Language
Code
Sentences
Malagasy
mlg
108,333… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/african-transcribed-speech-monolingual.nuer_greetings_monolingual_pairs_eng_nuer_dinkanuer_greetings_monolingual
License
This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
You are free to:
Share — copy and redistribute the dataset in any medium or format.
Adapt — remix, transform, and build upon the dataset for any purpose, including commercially.
Under the following terms:
Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made.
For more information, see the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/nuer_greetings_monolingual.manipuri-monolingual-corpus
Manipuri Monolingual Corpus
About
This Manipuri Monolingual corpus contains an expanded monolingual corpus for Manipuri in the following paper. It has been compiled from publicly available texts on the internet in the open domain.
Dataset Statistics
Set 1 Dataset contains approx. 11 million words.
Set 2 Dataset contains approx. 19 million words.
Set 3 Dataset contains approx. 76 million words.
Dataset Quality
Set 1 is of high quality. Set 3 is of low… See the full description on the dataset page: https://huggingface.co/datasets/joyson117/manipuri-monolingual-corpus.south_sudan_tourism_english_monolingual
