stevenyuan666/fineweb-edu-2017-qwen2-7b
FineWeb-Edu 2017 (~100B-token subset) with Qwen2-7B token counts A ~100B-token subset of FineWeb-Edu 2017, prepared for continued pretraining, with token counts computed by a pinned Qwen2-7B tokenizer. This dataset is a selected subset, not the complete 2017 crawl year. 2017 contains about 168B Qwen2-7B tokens, above the 100B target, so it was shuffled and subsetted: data/train/ holds 101,840,059 documents and 100,000,020,347 tokens, which is 59.29% of the 171,755,787 documents… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2017-qwen2-7b.
Clarify that 2017 is the selected ~100B-token subset, not the whole year
Publish full 2017 year: expose data/train, final totals, verification
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add 100-document 2017 validation sample with selection method and verification
initial commit
