CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes730 downloads4y agoHugging Face02indonesian-nlp /mc4-idA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.texttext-generation1M<n<10M14 likes557 downloads4y agoHugging Face03thegoodfellas /mc4-pt-cleaned Description This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4 Clean procedure We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.textfill-mask100M<n<1B4 likes342 downloads3y agoHugging Face04zicsx /mC4-hindi Dataset Card for "mC4-hindi" This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts. This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.texttext-generation10M<n<100M0 likes338 downloads3y agoHugging Face05yhavinga /mc4_nl_cleanedgated Dataset Card for Clean Dutch mC4 Dataset Summary A cleaned version (151GB) of the Dutch part (277GB) of the C4 multilingual dataset (mC4). Based on the Common Crawl dataset. The original version was prepared by AllenAI, hosted at the address https://huggingface.co/datasets/allenai/c4. Preprocessing The Dutch portion of mC4 was cleaned in a similar fashion as the English cleaned C4 version. See GitLab for details. In summary, the preprocessing procedure… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/mc4_nl_cleaned.texttext-generation100M<n<1B16 likes248 downloads1y agoHugging Face06Finnish-NLP /mc4_fi_cleaned Dataset Card for mC4 Finnish Cleaned Dataset Summary mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split. Supported Tasks and Leaderboards mC4 Finnish is mainly intended to pretrain Finnish language models and word representations. Languages Finnish Dataset Structure Data Instances [Needs More Information] Data Fields The data have several fields: url: url of the source as a string text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.texttext-generation10M<n<100M4 likes214 downloads4y agoHugging Face07jiviteshjn /mc4-zh-idiom-cpt mC4 zh — Idiom-Tagged Continued-Pretraining Corpus A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in figurative language. Each document is natural web text (from the C4/mC4 zh subset) containing at least one culturally meaningful chengyu, with an appended knowledge block that lists every matched idiom together with its figurative meaning(s) and classical source citation. Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.tabulartext-generation1M<n<10M0 likes59 downloads2mo agoHugging Face08nhagar /mc4-es-sampled_urls Dataset Card for mc4-es-sampled_urls This dataset provides the URLs and top-level domains associated with training records in bertin-project/mc4-es-sampled. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/mc4-es-sampled_urls.texttext-generation100M<n<1B0 likes7 downloads1y agoHugging Face09nhagar /clean_mc4_it_urls Dataset Card for clean_mc4_it_urls This dataset provides the URLs and top-level domains associated with training records in gsarti/clean_mc4_it. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/clean_mc4_it_urls.texttext-generation100M<n<1B0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.