codelion/finepdfs-1B
Sampling Methodology This dataset was created using reservoir sampling, a statistically unbiased random sampling algorithm that guarantees each sample from the source dataset has an equal probability of being included. This ensures the 1B token sample is representative of the full dataset's characteristics. Source Dataset: HuggingFaceFW/finepdfs Sample Size: 1B tokens Content: High-quality textbook-style pdfs Reservoir sampling enables rapid experimentation and ablation studies… See the full description on the dataset page: https://huggingface.co/datasets/codelion/finepdfs-1B.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face