CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MaLA-LM /mala-monolingual-filter MaLA Corpus: Massive Language Adaptation Corpus This is a cleaned version with some necessary data cleaning. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-filter.text-generation3 likes5.3k downloads2mo agoHugging Face02MaLA-LM /mala-monolingual-integration MaLA Corpus: Massive Language Adaptation Corpus This is the noisy version that integrates texts from different sources. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.texttext-generation1B<n<10B2 likes3.3k downloads2mo agoHugging Face03dsfsi /vukuzenzele-monolingual The Vuk'uzenzele South African Multilingual Corpus Give Feedback 📑: DSFSI Resource Feedback Form About Dataset The dataset was obtained from the South African government magazine Vuk'uzenzele, created by the Government Communication and Information System (GCIS). The original raw PDFs were obtatined from the Vuk'uzenzele website. The datasets contain government magazine editions in 11 languages, namely: Language Code Language Code English (eng) Sepedi (nso)… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/vukuzenzele-monolingual.texttranslation1K<n<10K4 likes3.3k downloads3y agoHugging Face04MaLA-LM /mala-monolingual-dedup MaLA Corpus: Massive Language Adaptation Corpus This is a deduplicated version after minhash and exact hash deduplication. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-dedup.text-generation2 likes2.8k downloads2mo agoHugging Face05MaLA-LM /mala-monolingual-split MaLA Corpus: Massive Language Adaptation Corpus This version contains train and validation splits. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.texttext-generation100M<n<1B4 likes2.6k downloads2mo agoHugging Face06catherinearnett /monolingual-tokenizer-dataTodo: add language to metadata cite source and explain sampling text100M<n<1B1 likes1.4k downloads1y agoHugging Face07Llamacha /monolingual-quechua-iic Dataset Card for Monolingual-Quechua-IIC Dataset Summary We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers. Supported Tasks and Leaderboards More Information Needed Languages Southern… See the full description on the dataset page: https://huggingface.co/datasets/Llamacha/monolingual-quechua-iic.textfill-mask100K<n<1M4 likes505 downloads4y agoHugging Face08spsither /tibetan_monolingual_A_merged_123_linestext100M<n<1B0 likes471 downloads2y agoHugging Face09LianHong /zomi-monolingual-corpus Zomi Monolingual Corpus v1.0 The Zomi Monolingual Corpus v1.0 contains 363,401 cleaned, deduplicated, reviewed, and permission-approved Zomi sentences. Zomi is represented with the ISO 639-3 language code ctd (Tedim Chin). Quick start from datasets import load_dataset dataset = load_dataset("LianHong/zomi-monolingual-corpus", split="train") print(dataset.num_rows) # 363401 print(dataset[0]["zomi_text"]) Data fields Field Type Description… See the full description on the dataset page: https://huggingface.co/datasets/LianHong/zomi-monolingual-corpus.texttext-generation100K<n<1M0 likes425 downloads26d agoHugging Face10nthakur /swim-ir-monolingual Dataset Card for SWIM-IR (Monolingual) This is the monolingual subset of the SWIM-IR dataset, where the query generated and the passage are both in the same language. A few remaining languages will be added in the upcoming v2 version of SWIM-IR. The dataset is available as CC-BY-SA 4.0. For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website. What is SWIM-IR? SWIM-IR dataset is a synthetic multilingual retrieval dataset… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/swim-ir-monolingual.texttext-retrieval1M<n<10M10 likes348 downloads2y agoHugging Face11spsither /tibetan_monolingual_Atext100M<n<1B1 likes340 downloads2y agoHugging Face12DigitalUmuganda /monolingual_machine_translation_datatext100K<n<1M0 likes337 downloads3y agoHugging Face13spsither /tibetan_monolingual_A_merged_135_linestext100M<n<1B0 likes320 downloads2y agoHugging Face14spsither /tibetan_monolingual_A_filteredtabular100M<n<1B0 likes247 downloads2y agoHugging Face15rewicks /ParaDocs-Monolingual0 likes225 downloads2y agoHugging Face16SimbaMaw1547 /south-african-monolingual-corpora-jsonl South African Languages Pretraining Dataset This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections. The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity Languages Included Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.text1M<n<10M0 likes209 downloads1y agoHugging Face17sanderland /monolingual-tokenizer-dataTodo: add language to metadata cite source and explain sampling text100M<n<1B0 likes208 downloads1y agoHugging Face18chronbmm /sanskrit-monolingual-pretraining-corrupted Dataset Card for "sanskrit-monolingual-pretraining-corrupted" More Information needed 10M<n<100M0 likes195 downloads3y agoHugging Face19michsethowusu /african-bible-monolingualtext1M<n<10M0 likes192 downloads3mo agoHugging Face20spsither /tibetan_monolingual_gold_merged_135_linestext10M<n<100M0 likes180 downloads2y agoHugging Face21ignatius /igbo_monolingualA dataset is a collection of Monolingual Igbo sentences.text-generation1K<n<10K2 likes169 downloads3y agoHugging Face22spsither /tibetan_monolingual_goldtext100M<n<1B0 likes152 downloads2y agoHugging Face23FrancophonIA /COVID-19_EU_presscorner_monolingual [!NOTE] Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19444 Description This collection of monolingual documents was generated from content available at https://ec.europa.eu/commission/presscorner. It includes 1354 documents in total in the following languages: EN 335 DE 266 ES 115 FR 123 IT 273 EL 120 SV 122 Citation EU presscorner monolingual collections of COVID-19 related documents. (2020, June 15). Version 1.0. [Dataset (Text corpus)].… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19_EU_presscorner_monolingual.translation0 likes144 downloads1y agoHugging Face24spsither /tibetan_monolingual_S_cleaned_train_test_tokenized10M<n<100M0 likes115 downloads2y agoHugging Face25failed09 /bashkir-wikipedia-monolingual Bashkir Wikipedia Monolingual Corpus Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer training and linguistic research. Overview Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801), cleaned and filtered with automated language identification. The cleaned configuration is the recommended default for language modelling, tokenization and linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.texttext-generation1M<n<10M0 likes104 downloads5d agoHugging Face26mbazaNLP /kinyarwanda_monolingual_v01.0 !!! PLEASE USE mbazaNLP/kinyarwanda_monolingual_v01.1 !!! !!! This version contains several duplicates and few non-kinyarwanda documents Dataset Summary The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts. This dataset contains 78k documents, totalling about 25 million words, and includes diverse content types such… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.0.tabulartext-generation10K<n<100K0 likes94 downloads2y agoHugging Face27Nikhil8076 /nepali-monolingual-dataset0 likes89 downloads1mo agoHugging Face28Nikhil8076 /telugu-monolingual-datasettext1M<n<10M0 likes87 downloads1mo agoHugging Face29spsither /tibetan_monolingual_Stext100M<n<1B0 likes84 downloads2y agoHugging Face30spsither /tibetan_monolingual_S_cleantext10M<n<100M0 likes82 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.