CoolFace
Datasetpublic

codelion/finepdfs-1B

Sampling Methodology This dataset was created using reservoir sampling, a statistically unbiased random sampling algorithm that guarantees each sample from the source dataset has an equal probability of being included. This ensures the 1B token sample is representative of the full dataset's characteristics. Source Dataset: HuggingFaceFW/finepdfs Sample Size: 1B tokens Content: High-quality textbook-style pdfs Reservoir sampling enables rapid experimentation and ablation studies… See the full description on the dataset page: https://huggingface.co/datasets/codelion/finepdfs-1B.

sourceHugging Faceupdated 11mo agoView on Hugging Face
4likes140downloads
3 commits on main
6fbca6d11mo ago

Add citation and methodology info

codelion
6a9e6201y ago

Upload dataset

codelion
93319e41y ago

initial commit

codelion