datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb_edu_100BT-shuffled
FineWeb-Edu 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.FineWeb-Edu-10B-Shuffleddclm-baseline-1.0-llama3-tokenized-shuffled
!! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !!
DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context)
This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines.
The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.dclm_100BT-shuffled
DCLM 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/dclm_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.dclm-baseline-1.0-llama3-tokenized-shuffled-524K
!! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !!
DCLM-Baseline Pretokenized (LLaMA 3.1, 524288 context)
This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines.
The original… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled-524K.MidTool-Mix-shuffledshuffled all sets
MidTool-Mix
A 20.3B-token mid-training corpus for agentic tool use. It pairs filtered web, PDF, and code sources with synthesized agent supervision, and is designed to teach models to recognize tool affordances, ground arguments from context, compose tool-call workflows, and recover from incomplete information — before any post-training.
Mid-training Qwen3-4B-Base / Qwen3-8B-Base on MidTool-Mix improves downstream tool use under both SFT and RL on BFCLv3… See the full description on the dataset page: https://huggingface.co/datasets/RedMod/MidTool-Mix-shuffled.MultiLegalPile_Wikipedia_Shuffledtok-corpus-shuffled
Dataset Card for "tok-corpus-shuffled"
This is the dataset used to fit custom tokenizers. Goal is to have a ytokenizer that is good for French, Englsih and Code.
The dataset uploaded is shuffled to facilitate subsampling it for tokenizer training.
French
Dataset({
features: ['id', 'text', 'dataset_id'],
num_rows: 16881941
})
Code
Dataset({
features: ['id', 'text', 'dataset_id'],
num_rows: 6338566
})
English
Dataset({
features: ['text', 'id', 'dataset_id']… See the full description on the dataset page: https://huggingface.co/datasets/manu/tok-corpus-shuffled.finepdfs_edu_100BT-shuffled
FinePDFs-Edu 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_edu_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as finepdfs_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_100BT-shuffled.fineweb_100BT-shuffled
FineWeb 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_100BT-shuffled.kl3m-data-sample-004-shuffled
KL3M Data Sample 004 (Shuffled)
This dataset contains a shuffled sample of 10 million examples from the KL3M Data Project, an initiative by the ALEA Institute providing copyright-clean training resources for large language models across legal, regulatory, and government domains.
The KL3M Data Project encompasses approximately 28 TB of compressed documents from authoritative sources including court opinions, government regulatory materials, corporate filings, intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-sample-004-shuffled.nemo23_shuffledfinepdfs_100BT-shuffled
FinePDFs 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_100BT-shuffled.TCGA-12K-parquet-shuffled
TCGA-12K Parquet (Shuffled)
Attribution
This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.fineweb-edu-dedup-shuffled
FineWeb-Edu-Dedup (Globally Shuffled)
A uniformly shuffled version of the FineWeb-Edu-Dedup subset from SmolLM-Corpus by HuggingFace.
Source Data
This dataset is derived from HuggingFaceTB/smollm-corpus, specifically the fineweb-edu-dedup subset. That subset is itself derived from FineWeb-Edu, a filtered and deduplicated extract of Common Crawl selected for educational content quality.
Property
Value
Source dataset
HuggingFaceTB/smollm-corpus
Source subset… See the full description on the dataset page: https://huggingface.co/datasets/JonathanMiddleton/fineweb-edu-dedup-shuffled.mmlu_shuffledverifiable-coding-problems-python_decontaminated-tested-shufflednews-ka-shuffled-ELECTRAidl-wds-pdfs-images-shuffledsnli-ve-shuffledShuffledDataset2chunked-shuffled-wikipedia20220301en-bookcorpusopen
Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The order of the items in this dataset has been shuffled,
meaning you don't have to use dataset.shuffle,
which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.smiles_transformer_shuffledfineweb-edu-shuffled
FineWeb-EDU Shuffled
Pre-shuffled versions of HuggingFaceFW/fineweb-edu.
Configs
Config
Shards
~Rows
Description
sample-100BT
~1800
~96M
100B token sample, shuffled
sample-350BT
~1340
~335M
350B token sample, shuffled, deduplicated against val
val
~18
~4.5M
Validation set (held out from 100BT)
Usage
from datasets import load_dataset
# Load 100B token training set
ds_100b = load_dataset("SandyResearch/fineweb-edu-shuffled"… See the full description on the dataset page: https://huggingface.co/datasets/SandyResearch/fineweb-edu-shuffled.fineweb-edu-sample-10BT-shuffled
📚 FineWeb-Edu (Shuffled)
The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves.
This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu.
Shuffling was performed using the following script:
import datasets
data = datasets.load_dataset(
"HuggingFaceFW/fineweb-edu",
"sample-10BT",
split="train",
streaming=False,
)
data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.massive_kb_v1_shuffled
MassIVE-KB v1 — globally-shuffled parquet
Tandem mass spectra from the MassIVE-KB v1 peptide
spectral library, repackaged as globally-shuffled, zstd-compressed Parquet shards
for streaming model training.
Why this repackaging
The source MassIVE-KB MGF is ordered by source raw file, with the same peptide
repeated in consecutive entries. Feeding that order to training leaves long runs
of correlated spectra inside any locally-shuffled read window, which produces
loss… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/massive_kb_v1_shuffled.00_filtered_shufflednews-ka-shuffled-DISTILBERT
