datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb_edu_100BT-shuffled
FineWeb-Edu 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.dclm_100BT-shuffled
DCLM 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/dclm_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.tok-corpus-shuffled
Dataset Card for "tok-corpus-shuffled"
This is the dataset used to fit custom tokenizers. Goal is to have a ytokenizer that is good for French, Englsih and Code.
The dataset uploaded is shuffled to facilitate subsampling it for tokenizer training.
French
Dataset({
features: ['id', 'text', 'dataset_id'],
num_rows: 16881941
})
Code
Dataset({
features: ['id', 'text', 'dataset_id'],
num_rows: 6338566
})
English
Dataset({
features: ['text', 'id', 'dataset_id']… See the full description on the dataset page: https://huggingface.co/datasets/manu/tok-corpus-shuffled.finepdfs_edu_100BT-shuffled
FinePDFs-Edu 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_edu_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as finepdfs_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_100BT-shuffled.fineweb_100BT-shuffled
FineWeb 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_100BT-shuffled.kl3m-data-sample-004-shuffled
KL3M Data Sample 004 (Shuffled)
This dataset contains a shuffled sample of 10 million examples from the KL3M Data Project, an initiative by the ALEA Institute providing copyright-clean training resources for large language models across legal, regulatory, and government domains.
The KL3M Data Project encompasses approximately 28 TB of compressed documents from authoritative sources including court opinions, government regulatory materials, corporate filings, intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-sample-004-shuffled.finepdfs_100BT-shuffled
FinePDFs 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_100BT-shuffled.TCGA-12K-parquet-shuffled
TCGA-12K Parquet (Shuffled)
Attribution
This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.fineweb-edu-dedup-shuffled
FineWeb-Edu-Dedup (Globally Shuffled)
A uniformly shuffled version of the FineWeb-Edu-Dedup subset from SmolLM-Corpus by HuggingFace.
Source Data
This dataset is derived from HuggingFaceTB/smollm-corpus, specifically the fineweb-edu-dedup subset. That subset is itself derived from FineWeb-Edu, a filtered and deduplicated extract of Common Crawl selected for educational content quality.
Property
Value
Source dataset
HuggingFaceTB/smollm-corpus
Source subset… See the full description on the dataset page: https://huggingface.co/datasets/JonathanMiddleton/fineweb-edu-dedup-shuffled.mmlu_shuffledverifiable-coding-problems-python_decontaminated-tested-shufflednews-ka-shuffled-ELECTRAmassive_kb_v1_shuffled
MassIVE-KB v1 — globally-shuffled parquet
Tandem mass spectra from the MassIVE-KB v1 peptide
spectral library, repackaged as globally-shuffled, zstd-compressed Parquet shards
for streaming model training.
Why this repackaging
The source MassIVE-KB MGF is ordered by source raw file, with the same peptide
repeated in consecutive entries. Feeding that order to training leaves long runs
of correlated spectra inside any locally-shuffled read window, which produces
loss… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/massive_kb_v1_shuffled.snli-ve-shuffledShuffledDataset2smiles_transformer_shuffledidl-wds-pdfs-images-shuffledfineweb-edu-shuffled
FineWeb-EDU Shuffled
Pre-shuffled versions of HuggingFaceFW/fineweb-edu.
Configs
Config
Shards
~Rows
Description
sample-100BT
~1800
~96M
100B token sample, shuffled
sample-350BT
~1340
~335M
350B token sample, shuffled, deduplicated against val
val
~18
~4.5M
Validation set (held out from 100BT)
Usage
from datasets import load_dataset
# Load 100B token training set
ds_100b = load_dataset("SandyResearch/fineweb-edu-shuffled"… See the full description on the dataset page: https://huggingface.co/datasets/SandyResearch/fineweb-edu-shuffled.fineweb-edu-sample-10BT-shuffled
📚 FineWeb-Edu (Shuffled)
The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves.
This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu.
Shuffling was performed using the following script:
import datasets
data = datasets.load_dataset(
"HuggingFaceFW/fineweb-edu",
"sample-10BT",
split="train",
streaming=False,
)
data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.chunked-shuffled-wikipedia20220301en-bookcorpusopen
Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The order of the items in this dataset has been shuffled,
meaning you don't have to use dataset.shuffle,
which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.00_filtered_shufflednews-ka-shuffled-DISTILBERT153-angiosperm-species-32k-sequences-shuffledfineweb-edu-dedup-10b-30gram-shuffledMAPS-ClinVar-VKS-Embeddings-L80-shuffled
MAPS ClinVar/VKS ESM-C layer-80 difference fields, shuffled layout
The same 200,913 rows as
ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80, rewritten in one
random order so that a prefix is a sample.
The source repo is sorted by protein accession. That makes the first N rows a
block of related proteins rather than a sample of the pool, so any subset that
is actually representative requires reading all 154.01 GB and
selecting rows afterwards. Here the rows are stored shuffled, so… See the full description on the dataset page: https://huggingface.co/datasets/ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80-shuffled.dclm-16m-shuffledarctic-full-combined-shuffledfineweb-edu-Llama-3.2-Instruct-Shuffled
