datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
mc4-ja
Dataset Card for "mc4-ja"
More Information needed
clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual
colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning
detailed in the repository README file.mc4_3.1.0_fi_cleaned
Dataset Card for "mc4_3.1.0_fi_cleaned"
More Information needed
mc4-ja-filter-ja-normal
Dataset Card for "mc4-ja-filter-ja-normal"
More Information needed
mc4A colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI.zh-en-mc4mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.mC4-Hindi-Cleaned
Dataset Card for "mC4-Hindi-Cleaned"
More Information needed
mc4-pt
MC4-PT
MC4-PT is the is the portuguese subset from MC4.
MC4 is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org".
This is the raw version. Deduplicated version is available here.
temiz-mC4
Dataset Card for Temiz mC4
Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
split
num instances
size
num of words
train
76.432.893
168GB
21.06B
Total
76.432.893
168GB
21.06B
This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering and… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.mc4_es_cl
Dataset Card for "mc4_es_cl"
More Information needed
mc4-idA thoroughly cleaned version of the Italian portion of the multilingual
colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning
detailed in the repository README file.mC4-TESTmc4_und_idfiltered,deduplication MC4-ID from MC4 part undfined
mc4_legal
Dataset Card for MC4_Legal: A Corpus Covering the Legal Part of MC4 for European Languages
Dataset Summary
This dataset contains large text resources (~133GB in total) from mc4 filtered for legal data that can be used for pretraining language models.
Use the dataset like this:
from datasets import load_dataset
dataset = load_dataset("joelito/mc4_legal", "de", split='train', streaming=True)
Supported Tasks and Leaderboards
The dataset supports the task of… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/mc4_legal.mc4-fractionViet-Font-mc4-Textmc4-pt-cleaned
Description
This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4
Clean procedure
We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git
The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a
pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.mC4-hindi
Dataset Card for "mC4-hindi"
This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts.
This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.legal-mc4Legal-MC4: A Corpus Covering the Legal Part of MC4 for European Languagesenglish-mc4
Dataset Card for "english-mc4"
More Information needed
mc4_nl_cleaned
Dataset Card for Clean Dutch mC4
Dataset Summary
A cleaned version (151GB) of the Dutch part (277GB) of the C4 multilingual dataset (mC4).
Based on the Common Crawl dataset.
The original version was prepared by AllenAI, hosted at the address https://huggingface.co/datasets/allenai/c4.
Preprocessing
The Dutch portion of mC4 was cleaned in a similar fashion as the English cleaned C4 version.
See GitLab for details.
In summary, the preprocessing procedure… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/mc4_nl_cleaned.mc4-mytedetails_BramVanroy__llama2-13b-ft-mc4_nl_cleaned_tiny
Dataset Card for Evaluation run of BramVanroy/llama2-13b-ft-mc4_nl_cleaned_tiny
Dataset Summary
Dataset automatically created during the evaluation run of model BramVanroy/llama2-13b-ft-mc4_nl_cleaned_tiny on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_BramVanroy__llama2-13b-ft-mc4_nl_cleaned_tiny.mc4_fi_cleaned
Dataset Card for mC4 Finnish Cleaned
Dataset Summary
mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split.
Supported Tasks and Leaderboards
mC4 Finnish is mainly intended to pretrain Finnish language models and word representations.
Languages
Finnish
Dataset Structure
Data Instances
[Needs More Information]
Data Fields
The data have several fields:
url: url of the source as a string
text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.mc4-samplingA sampling-enabled version of mC4, the colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is a version of the processed version of Google's mC4 dataset by AllenAI, in which sampling methods are implemented to perform on the fly.mc4_validationA colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI.register_mc4mc4-filter
