CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /MADLAD-400 MADLAD-400 Dataset and Introduction MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is a document-level multilingual dataset based on Common Crawl, covering 419 languages in total. This uses all snapshots of CommonCrawl available as of August 1, 2022. The primary advantage of this dataset over similar datasets is that it is more multilingual (419 languages), it is audited and more highly filtered, and it is document-level. The main… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MADLAD-400.text-generationn>1T173 likes31k downloads2y agoHugging Face02malaysia-ai /mosaic-madlad-400-ms Mosaic format for extra dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms load it, from streaming import LocalDataset import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.textn<1K0 likes1.7k downloads3y agoHugging Face03Hemanth-thunder /tamil-madlad-400text1M<n<10M2 likes567 downloads3y agoHugging Face04ymoslem /MADLAD-Arabic-Cleantext10M<n<100M0 likes457 downloads2y agoHugging Face05ymoslem /MADLAD-Arabic-Clean-Flattenedtext100M<n<1B0 likes455 downloads2y agoHugging Face06quickmt /madlad400-en-backtranslated-ar madlad400 en Sample Translated into ar This dataset is a subset of MADLAD-400 translated from en into ar by the quickmt/quickmt-en-ar model (beam size 4) intended to be used for training translation models from ar into en. text10M<n<100M0 likes431 downloads9mo agoHugging Face07nhagar /madlad-400_urls_noisy Dataset Card for madlad-400_urls_noisy This dataset provides the URLs and top-level domains associated with training records in allenai/MADLAD-400 (noisy variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/madlad-400_urls_noisy.texttext-generation1B<n<10B0 likes177 downloads1y agoHugging Face08nhagar /madlad-400_urls_clean Dataset Card for madlad-400_urls_clean This dataset provides the URLs and top-level domains associated with training records in allenai/MADLAD-400 (clean variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/madlad-400_urls_clean.texttext-generation1B<n<10B0 likes168 downloads1y agoHugging Face09ccde /MADLAD-400-First2Mtext100K<n<1M0 likes163 downloads4mo agoHugging Face10Symato /madlad-400_vi MADLAD-400 Dataset and Introduction MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is a document-level multilingual dataset based on Common Crawl, covering 419 languages in total. This uses all snapshots of CommonCrawl available as of August 1, 2022. The primary advantage of this dataset over similar datasets is that it is more multilingual (419 languages), it is audited and more highly filtered, and it is document-level. The main disadvantage… See the full description on the dataset page: https://huggingface.co/datasets/Symato/madlad-400_vi.texttext-generation10M<n<100M1 likes162 downloads2y agoHugging Face11polyglots /MADLAD_CulturaX_cleanedtext10M<n<100M21 likes106 downloads2y agoHugging Face12mahan13900 /MADLAD-400_Persian-Clain.Finaly MADLAD-400 Persian Cleaned (Final) این دیتاست نسخه تمیزسازی‌شده و بهینه‌سازی‌شده از زیرمجموعه زبان فارسی دیتاست MADLAD-400 است. مشخصات دیتاست: زبان: فارسی (fa) تعداد کل رکوردها: 24,109,659 فرمت ذخیره‌سازی: Apache Parquet (Snappy) فیلد داده: text (متن تمیز فارسی) نحوه استفاده: from datasets import load_dataset dataset = load_dataset("mahan13900/MADLAD-400_Persian-Clain.Finaly") print(dataset["train"][0]) texttext-generation10M<n<100M1 likes104 downloads15d agoHugging Face13mahan13900 /madlad-fa-cleaned-lvl30 likes90 downloads17d agoHugging Face14realtmxi /MADLAD_CultureX_cleaned0 likes84 downloads1y agoHugging Face15quickmt /madlad400-en-backtranslated-ja madlad400 en Sample Translated into ja This dataset is a subset of MADLAD-400 translated from en into ja by the quickmt/quickmt-en-ja model (beam size 4) intended to be used for training translation models from ja into en. text10M<n<100M0 likes81 downloads9mo agoHugging Face16suwaimyo /madlad400-tet-classification MADLAD400_tet_Classification Deduplicated copy of kornwtp/madlad400-tet-classification. Splits split rows train 8,071 text1K<n<10K0 likes78 downloads28d agoHugging Face17neody /madlad-400-ja-cleanedtext1M<n<10M0 likes77 downloads2y agoHugging Face18puttatidam /madlad400-tet-bitextmining madlad400-tet-bitextmining Deduplicated copy of kornwtp/madlad400-tet-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/madlad400-tet-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/madlad400-tet-bitextmining.text10K<n<100K0 likes75 downloads10d agoHugging Face19Minuri /madlad_cleaned_version Sinhala Cleaned Sentences - MADLAD-400 Cleaned and deduplicated Sinhala sentences derived from Minuri/sinhala-corpus-madlad400, produced through a multi-stage cleaning pipeline. This repo was used as pipeline storage across cleaning stages, with the final output being stage10_final_corpus_deduped.csv. Final Output File Rows Description stage10_final_corpus_deduped.csv 5,033,732 Final cleaned and deduplicated sentences Dataset Structure (final… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/madlad_cleaned_version.text-generation1M<n<10M0 likes69 downloads6mo agoHugging Face20SEACrowd /sea_madladSEA MADLAD is a subset of MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level), which is a document-level multilingual dataset based on Common Crawl. SEA MADLAD only filters the language of the "clean" subset, which covers 36 languages indigenous to SEA from 419 languages in total. As a result, some of SEA lang codes aren't available in this version because those belongs to the languages whose decision was to "remove from its clean version" based on MADLAD auditing… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/sea_madlad.0 likes54 downloads2y agoHugging Face21quickmt /madlad400-en-backtranslated-fr madlad400 en Sample Translated into fr This dataset is a subset of MADLAD-400 translated from en into fr by the quickmt/quickmt-en-fr model (beam size 4) intended to be used for training translation models from fr into en. texttranslation10M<n<100M0 likes49 downloads8mo agoHugging Face22ccde /madlad-subset-latin-relatedtext10K<n<100K0 likes45 downloads2mo agoHugging Face23kornwtp /madlad400-tet-classificationtext1K<n<10K0 likes42 downloads2y agoHugging Face24quickmt /madlad400-en-backtranslated-zh madlad400 en Sample Translated into zh This dataset is a subset of MADLAD-400 translated from en into zh by the quickmt/quickmt-en-zh model (beam size 4) intended to be used for training translation models from zh into en. text10M<n<100M0 likes38 downloads10mo agoHugging Face25udmurtNLP /madlad-400-udmurt Usage madlad-400-udmurt from datasets import load_dataset dataset = load_dataset("udmurtNLP/madlad-400-udmurt") text100K<n<1M0 likes37 downloads3y agoHugging Face26Minuri /sinhala-corpus-madlad400 Sinhala Raw Sentences - MADLAD-400 Raw Sinhala sentences extracted and sentence-split from the allenai/MADLAD-400 dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (madlad) Split Rows train 7,281,026 Pipeline Position allenai/MADLAD-400 → this repo → Minuri/madlad_cleaned_version →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-madlad400.texttext-generation1M<n<10M0 likes37 downloads6mo agoHugging Face27malaysia-ai /madlad-400-ms0 likes35 downloads3y agoHugging Face28quickmt /madlad400-en-backtranslated-fa madlad400 en Sample Translated into fa This dataset is a subset of MADLAD-400 translated from en into fa by the quickmt/quickmt-en-fa model (beam size 4) intended to be used for training translation models from fa into en. texttranslation10M<n<100M0 likes34 downloads6mo agoHugging Face29ccde /madlad-subset-spaintext10K<n<100K0 likes32 downloads2mo agoHugging Face30ccde /madlad-subset-pngtext1K<n<10K0 likes30 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.