datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MADLAD-400
MADLAD-400
Dataset and Introduction
MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is
a document-level multilingual dataset based on Common Crawl, covering 419
languages in total. This uses all snapshots of CommonCrawl available as of August
1, 2022. The primary advantage of this dataset over similar datasets is that it
is more multilingual (419 languages), it is audited and more highly filtered,
and it is document-level. The main… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MADLAD-400.mosaic-madlad-400-ms
Mosaic format for extra dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms
load it,
from streaming import LocalDataset
import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.tamil-madlad-400MADLAD-Arabic-CleanMADLAD-Arabic-Clean-Flattenedmadlad400-en-backtranslated-ar
madlad400 en Sample Translated into ar
This dataset is a subset of MADLAD-400 translated from en into ar by the quickmt/quickmt-en-ar model (beam size 4) intended to be used for training translation models from ar into en.
madlad-400_urls_noisy
Dataset Card for madlad-400_urls_noisy
This dataset provides the URLs and top-level domains associated with training records in allenai/MADLAD-400 (noisy variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/madlad-400_urls_noisy.madlad-400_urls_clean
Dataset Card for madlad-400_urls_clean
This dataset provides the URLs and top-level domains associated with training records in allenai/MADLAD-400 (clean variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/madlad-400_urls_clean.MADLAD-400-First2Mmadlad-400_vi
MADLAD-400
Dataset and Introduction
MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is
a document-level multilingual dataset based on Common Crawl, covering 419
languages in total. This uses all snapshots of CommonCrawl available as of August
1, 2022. The primary advantage of this dataset over similar datasets is that it
is more multilingual (419 languages), it is audited and more highly filtered,
and it is document-level. The main disadvantage… See the full description on the dataset page: https://huggingface.co/datasets/Symato/madlad-400_vi.MADLAD_CulturaX_cleanedMADLAD-400_Persian-Clain.Finaly
MADLAD-400 Persian Cleaned (Final)
این دیتاست نسخه تمیزسازیشده و بهینهسازیشده از زیرمجموعه زبان فارسی دیتاست MADLAD-400 است.
مشخصات دیتاست:
زبان: فارسی (fa)
تعداد کل رکوردها: 24,109,659
فرمت ذخیرهسازی: Apache Parquet (Snappy)
فیلد داده: text (متن تمیز فارسی)
نحوه استفاده:
from datasets import load_dataset
dataset = load_dataset("mahan13900/MADLAD-400_Persian-Clain.Finaly")
print(dataset["train"][0])
madlad-fa-cleaned-lvl3MADLAD_CultureX_cleanedmadlad400-en-backtranslated-ja
madlad400 en Sample Translated into ja
This dataset is a subset of MADLAD-400 translated from en into ja by the quickmt/quickmt-en-ja model (beam size 4) intended to be used for training translation models from ja into en.
madlad400-tet-classification
MADLAD400_tet_Classification
Deduplicated copy of kornwtp/madlad400-tet-classification.
Splits
split
rows
train
8,071
madlad-400-ja-cleanedmadlad400-tet-bitextmining
madlad400-tet-bitextmining
Deduplicated copy of kornwtp/madlad400-tet-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/madlad400-tet-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/madlad400-tet-bitextmining.madlad_cleaned_version
Sinhala Cleaned Sentences - MADLAD-400
Cleaned and deduplicated Sinhala sentences derived from Minuri/sinhala-corpus-madlad400, produced through a multi-stage cleaning pipeline. This repo was used as pipeline storage across cleaning stages, with the final output being stage10_final_corpus_deduped.csv.
Final Output
File
Rows
Description
stage10_final_corpus_deduped.csv
5,033,732
Final cleaned and deduplicated sentences
Dataset Structure (final… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/madlad_cleaned_version.sea_madladSEA MADLAD is a subset of MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level), which is a document-level multilingual dataset based on Common Crawl.
SEA MADLAD only filters the language of the "clean" subset, which covers 36 languages indigenous to SEA from 419 languages in total.
As a result, some of SEA lang codes aren't available in this version because those belongs to the languages whose decision was to "remove from its clean version" based on MADLAD auditing… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/sea_madlad.madlad400-en-backtranslated-fr
madlad400 en Sample Translated into fr
This dataset is a subset of MADLAD-400 translated from en into fr by the quickmt/quickmt-en-fr model (beam size 4) intended to be used for training translation models from fr into en.
madlad-subset-latin-relatedmadlad400-tet-classificationmadlad400-en-backtranslated-zh
madlad400 en Sample Translated into zh
This dataset is a subset of MADLAD-400 translated from en into zh by the quickmt/quickmt-en-zh model (beam size 4) intended to be used for training translation models from zh into en.
madlad-400-udmurt
Usage madlad-400-udmurt
from datasets import load_dataset
dataset = load_dataset("udmurtNLP/madlad-400-udmurt")
sinhala-corpus-madlad400
Sinhala Raw Sentences - MADLAD-400
Raw Sinhala sentences extracted and sentence-split from the allenai/MADLAD-400 dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset.
Dataset Structure
Column
Description
text
Raw Sinhala sentence
source
Source identifier (madlad)
Split
Rows
train
7,281,026
Pipeline Position
allenai/MADLAD-400 → this repo → Minuri/madlad_cleaned_version →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-madlad400.madlad-400-msmadlad400-en-backtranslated-fa
madlad400 en Sample Translated into fa
This dataset is a subset of MADLAD-400 translated from en into fa by the quickmt/quickmt-en-fa model (beam size 4) intended to be used for training translation models from fa into en.
madlad-subset-spainmadlad-subset-png
