CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SimbaMaw1547 /south-african-monolingual-corpora-jsonl South African Languages Pretraining Dataset This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections. The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity Languages Included Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.text1M<n<10M0 likes196 downloads1y agoHugging Face02Nikhil8076 /telugu-monolingual-datasettext1M<n<10M0 likes91 downloads1mo agoHugging Face03AfriSpeech /african-transcribed-speech-monolingual African Transcribed Speech — Monolingual Sentence-level monolingual text for 15 African languages, derived from translated religious speech transcriptions. Intended as reference text for evaluating speech machine translation (e.g. BLEU scoring), and as a monolingual corpus for language modeling / tokenizer training. Each language is a separate subset — load with e.g. load_dataset("<repo>", "fat"). Coverage Language Code Sentences Malagasy mlg 108,333… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/african-transcribed-speech-monolingual.texttext-generation100K<n<1M0 likes38 downloads1mo agoHugging Face04NaathNLP /nuer_greetings_monolingual_pairs_eng_nuer_dinkatext10K<n<100K0 likes35 downloads1mo agoHugging Face05dayomtechnologies /nuer_greetings_monolingual License This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to: Share — copy and redistribute the dataset in any medium or format. Adapt — remix, transform, and build upon the dataset for any purpose, including commercially. Under the following terms: Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. For more information, see the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/nuer_greetings_monolingual.text100K<n<1M0 likes30 downloads1mo agoHugging Face06joyson117 /manipuri-monolingual-corpusgated Manipuri Monolingual Corpus About This Manipuri Monolingual corpus contains an expanded monolingual corpus for Manipuri in the following paper. It has been compiled from publicly available texts on the internet in the open domain. Dataset Statistics Set 1 Dataset contains approx. 11 million words. Set 2 Dataset contains approx. 19 million words. Set 3 Dataset contains approx. 76 million words. Dataset Quality Set 1 is of high quality. Set 3 is of low… See the full description on the dataset page: https://huggingface.co/datasets/joyson117/manipuri-monolingual-corpus.texttext-generation100K<n<1M1 likes14 downloads1y agoHugging Face07dayomtechnologies /south_sudan_tourism_english_monolingualtext100K<n<1M0 likes3 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.