CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes4.1k downloads3y agoHugging Face02izumi-lab /mc4-ja Dataset Card for "mc4-ja" More Information needed text10M<n<100M6 likes4k downloads3y agoHugging Face03gsarti /clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.text-generation100M<n<1B18 likes2.1k downloads2y agoHugging Face04Finnish-NLP /mc4_3.1.0_fi_cleaned Dataset Card for "mc4_3.1.0_fi_cleaned" More Information needed tabular10M<n<100M0 likes1.4k downloads3y agoHugging Face05izumi-lab /mc4-ja-filter-ja-normal Dataset Card for "mc4-ja-filter-ja-normal" More Information needed text10M<n<100M5 likes972 downloads3y agoHugging Face06legacy-datasets /mc4A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI.text-generationn<1K155 likes968 downloads3y agoHugging Face07blobba /zh-en-mc42 likes755 downloads4y agoHugging Face08bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes730 downloads4y agoHugging Face09zicsx /mC4-Hindi-Cleaned Dataset Card for "mC4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes691 downloads3y agoHugging Face10eduagarcia /mc4-pt MC4-PT MC4-PT is the is the portuguese subset from MC4. MC4 is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the raw version. Deduplicated version is available here. text100M<n<1B2 likes614 downloads3y agoHugging Face11turkish-nlp-suite /temiz-mC4 Dataset Card for Temiz mC4 Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. split num instances size num of words train 76.432.893 168GB 21.06B Total 76.432.893 168GB 21.06B This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering and… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.fill-mask2 likes580 downloads11mo agoHugging Face12jorgeortizfuentes /mc4_es_cl Dataset Card for "mc4_es_cl" More Information needed text1M<n<10M1 likes559 downloads4y agoHugging Face13indonesian-nlp /mc4-idA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.texttext-generation1M<n<10M14 likes557 downloads4y agoHugging Face14markus583 /mC4-TESTtext100M<n<1B0 likes468 downloads3y agoHugging Face15acul3 /mc4_und_idfiltered,deduplication MC4-ID from MC4 part undfined text1M<n<10M0 likes446 downloads2y agoHugging Face16joelniklaus /mc4_legal Dataset Card for MC4_Legal: A Corpus Covering the Legal Part of MC4 for European Languages Dataset Summary This dataset contains large text resources (~133GB in total) from mc4 filtered for legal data that can be used for pretraining language models. Use the dataset like this: from datasets import load_dataset dataset = load_dataset("joelito/mc4_legal", "de", split='train', streaming=True) Supported Tasks and Leaderboards The dataset supports the task of… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/mc4_legal.textfill-mask1M<n<10M7 likes416 downloads4y agoHugging Face17bowphs /mc4-fractiontext1M<n<10M0 likes416 downloads2y agoHugging Face185CD-AI /Viet-Font-mc4-Textimage10K<n<100K3 likes368 downloads2y agoHugging Face19thegoodfellas /mc4-pt-cleaned Description This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4 Clean procedure We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.textfill-mask100M<n<1B4 likes342 downloads3y agoHugging Face20zicsx /mC4-hindi Dataset Card for "mC4-hindi" This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts. This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.texttext-generation10M<n<100M0 likes338 downloads3y agoHugging Face21joelniklaus /legal-mc4Legal-MC4: A Corpus Covering the Legal Part of MC4 for European Languagestextfill-mask1M<n<10M14 likes334 downloads3y agoHugging Face22tiennv /english-mc4 Dataset Card for "english-mc4" More Information needed text10M<n<100M0 likes331 downloads3y agoHugging Face23yhavinga /mc4_nl_cleanedgated Dataset Card for Clean Dutch mC4 Dataset Summary A cleaned version (151GB) of the Dutch part (277GB) of the C4 multilingual dataset (mC4). Based on the Common Crawl dataset. The original version was prepared by AllenAI, hosted at the address https://huggingface.co/datasets/allenai/c4. Preprocessing The Dutch portion of mC4 was cleaned in a similar fashion as the English cleaned C4 version. See GitLab for details. In summary, the preprocessing procedure… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/mc4_nl_cleaned.texttext-generation100M<n<1B16 likes248 downloads1y agoHugging Face24lingvenvist /mc4-myte0 likes232 downloads4mo agoHugging Face25open-llm-leaderboard-old /details_BramVanroy__llama2-13b-ft-mc4_nl_cleaned_tiny Dataset Card for Evaluation run of BramVanroy/llama2-13b-ft-mc4_nl_cleaned_tiny Dataset Summary Dataset automatically created during the evaluation run of model BramVanroy/llama2-13b-ft-mc4_nl_cleaned_tiny on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_BramVanroy__llama2-13b-ft-mc4_nl_cleaned_tiny.0 likes231 downloads3y agoHugging Face26Finnish-NLP /mc4_fi_cleaned Dataset Card for mC4 Finnish Cleaned Dataset Summary mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split. Supported Tasks and Leaderboards mC4 Finnish is mainly intended to pretrain Finnish language models and word representations. Languages Finnish Dataset Structure Data Instances [Needs More Information] Data Fields The data have several fields: url: url of the source as a string text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.texttext-generation10M<n<100M4 likes214 downloads4y agoHugging Face27bertin-project /mc4-samplingA sampling-enabled version of mC4, the colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is a version of the processed version of Google's mC4 dataset by AllenAI, in which sampling methods are implemented to perform on the fly.text-generationn<1K13 likes206 downloads2y agoHugging Face28vocabtrimmer /mc4_validationA colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI.text1M<n<10M0 likes200 downloads4y agoHugging Face29TurkuNLP /register_mc4text1M<n<10M0 likes131 downloads5y agoHugging Face30Wikidepia /mc4-filter0 likes126 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.