CoolFace
Datasetpublic

HuggingFaceFW/finepdfs_100BT-shuffled

FinePDFs 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_100BT-shuffled.

sourceHugging Faceodc-byupdated 7mo agoView on Hugging Face
0likes617downloads
Dataset Card

FinePDFs 100BT (Shuffled)

A globally shuffled version of HuggingFaceFW/finepdfs_100BT.

Part of the Smol-Data collection — tried and tested mixes for strong pretraining.

Dataset Description

This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.

How It Was Created

The unshuffled dataset was loaded into memory, shuffled with dataset.shuffle(seed=42), and re-uploaded with 100 shards. See the smol_data.py script for details.

Usage

python
from datasets import load_dataset

ds = load_dataset("HuggingFaceFW/finepdfs_100BT-shuffled", split="train", streaming=True)
for sample in ds:
    print(sample["text"][:200])
    break

Citation

bibtex
@misc{niklaus2026smoldata,
      title={SmolData},
      author={Joel Niklaus and Hynek Kydl{\'\i}{\v{c}}ek},
      year={2026},
      publisher={Hugging Face},
      journal={Hugging Face repository},
      howpublished={\url{https://huggingface.co/collections/HuggingFaceFW/smol-data}}
}