HuggingFaceFW/finepdfs_100BT-shuffled
FinePDFs 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_100BT-shuffled.
FinePDFs 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with dataset.shuffle(seed=42), and re-uploaded with 100 shards. See the smol_data.py script for details.
Usage
from datasets import load_dataset
ds = load_dataset("HuggingFaceFW/finepdfs_100BT-shuffled", split="train", streaming=True)
for sample in ds:
print(sample["text"][:200])
breakCitation
@misc{niklaus2026smoldata,
title={SmolData},
author={Joel Niklaus and Hynek Kydl{\'\i}{\v{c}}ek},
year={2026},
publisher={Hugging Face},
journal={Hugging Face repository},
howpublished={\url{https://huggingface.co/collections/HuggingFaceFW/smol-data}}
}