CoolFace
Datasetpublic

stevenyuan666/fineweb-edu-2017-qwen2-7b

FineWeb-Edu 2017 (~100B-token subset) with Qwen2-7B token counts A ~100B-token subset of FineWeb-Edu 2017, prepared for continued pretraining, with token counts computed by a pinned Qwen2-7B tokenizer. This dataset is a selected subset, not the complete 2017 crawl year. 2017 contains about 168B Qwen2-7B tokens, above the 100B target, so it was shuffled and subsetted: data/train/ holds 101,840,059 documents and 100,000,020,347 tokens, which is 59.29% of the 171,755,787 documents… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2017-qwen2-7b.

sourceHugging Faceodc-byupdated 17d agoView on Hugging Face
0likes356downloads

stevenyuan666/fineweb-edu-2017-qwen2-7b · main · files are served by the source, never re-hosted here