CoolFace
Datasetpublic

codelion/finepdfs-100M

Sampling Methodology This dataset was created using reservoir sampling, a statistically unbiased random sampling algorithm that guarantees each sample from the source dataset has an equal probability of being included. This ensures the 100M token sample is representative of the full dataset's characteristics. Source Dataset: HuggingFaceFW/finepdfs Sample Size: 100M tokens Content: High-quality textbook-style pdfs Reservoir sampling enables rapid experimentation and ablation… See the full description on the dataset page: https://huggingface.co/datasets/codelion/finepdfs-100M.

sourceHugging Faceupdated 11mo agoView on Hugging Face
2likes17downloads
settings

This repository belongs to codelion on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefinepdfs-100M
visibilitypublic
licencenot set
gatedno
ownercodelion
Account settings
codelion/finepdfs-100M · CoolFace