mC4
BramVanroy_-_llama2-13b-ft-mc4_nl_cleaned_tiny-ggufmultilingual-pruning_-_pruned-aya-23-8b-wanda-0.5-unstructured-mc4-ar-0-ggufjina-text-matching-embeddings-v5-text-nano-retrieval-mc4-3fre-v3convbert-base-turkish-mc4-uncasedjina-embeddings-v5-text-nano-retrieval-df-3fre-mc4-80kt5-base-japanese-mC4-Wikipediaelectra-base-turkish-mc4-uncased-discriminatorconvbert-base-turkish-mc4-toxicity-uncased
Datasets
All datasets matching “mC4”mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
mc4-ja
Dataset Card for "mc4-ja"
More Information needed
clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual
colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning
detailed in the repository README file.mc4_3.1.0_fi_cleaned
Dataset Card for "mc4_3.1.0_fi_cleaned"
More Information needed
mc4-ja-filter-ja-normal
Dataset Card for "mc4-ja-filter-ja-normal"
More Information needed
mc4A colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI.
