Usmansafder/fineweb-edu-2014-qwen2
FineWeb-Edu 2014, Qwen2-7B token counts 86,901,732 documents, 92,348,618,822 Qwen2-7B tokens, prepared for continued pretraining as part of the FinMoE project. column type meaning date int32 the FineWeb year, 2014 text string document text, unmodified token_count int32 Qwen2-7B tokens in text Source HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 8 CommonCrawl dumps of 2014 (115 shards). Token… See the full description on the dataset page: https://huggingface.co/datasets/Usmansafder/fineweb-edu-2014-qwen2.
FineWeb-Edu 2014, Qwen2-7B token counts
86,901,732 documents, 92,348,618,822 Qwen2-7B tokens, prepared for continued pretraining as part of the FinMoE project.
Source
HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 8 CommonCrawl dumps of 2014 (115 shards).
Token counts
Qwen/Qwen2-7B at revision 453ed1575b739b5b03ce3758b23befdb0967f40e, counted on the raw text with add_special_tokens=False and truncation=False. FineWeb-Edu's own token_count column is GPT-2 based and was discarded, not copied.
Selection
The year holds 92.35 B Qwen2 tokens, below the 100B target, so every document is retained: nothing was shuffled away or subset. Each document is still assigned u, a uniform draw from random.Random(20140101 + shard_index) taken in row order, and is stored in ascending u order within its shard; the cutoff u < 1.0 therefore selects the whole year.
Full details, including package versions, are in processing_metadata.json.
