CoolFace
Datasetpublic

BoomQ/fineweb-edu-2016-qwen2

FineWeb-Edu 2016 / Qwen2 Completed 2016 crawl-year processing. Documents: 93,099,356. Actual recounted Qwen2 tokens: 99,541,846,844. Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9. The input inventory covers 9 crawl directories. date is the integer crawl year 2016, not an article publication date. Original text is preserved without cleaning, normalization, deduplication, truncation, or added formatting. Source token counts are not used.… See the full description on the dataset page: https://huggingface.co/datasets/BoomQ/fineweb-edu-2016-qwen2.

sourceHugging Faceodc-byupdated 16d agoView on Hugging Face
0likes105downloads
Dataset Card

FineWeb-Edu 2016 / Qwen2

Completed 2016 crawl-year processing.

Documents: 93,099,356. Actual recounted Qwen2 tokens: 99,541,846,844.

Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9. The input inventory covers 9 crawl directories. date is the integer crawl year 2016, not an article publication date. Original text is preserved without cleaning, normalization, deduplication, truncation, or added formatting. Source token counts are not used. Optional source/category columns are omitted.

Tokenizer: Qwen/Qwen2-7B, revision 453ed1575b739b5b03ce3758b23befdb0967f40e, use_fast=True, trust_remote_code=False; Transformers 5.0.0, tokenizers 0.22.2. Each raw text is tokenized with add_special_tokens=False, truncation=False, padding=False; attention masks and token type IDs are disabled. The saved count is the length of input_ids. Documents beyond the model context length remain intact.

Full provenance, exact counts, file checksums, and selection details are in manifest.json. The pinned input file inventory is in source_manifest.json.

The full year contained 93,099,356 documents and 99,541,846,844 actual Qwen2 tokens. Selection details:

json
{
  "hash_collision_tiebreak": "ascending canonical source index, then row index",
  "output_order": "canonical source order, then original row order; selected membership is shuffled",
  "prefix_rule": "nearest actual token sum including empty prefix; equal distance chooses smaller token total, then shorter prefix",
  "rule": "retain_all",
  "seed": 2016,
  "shuffle_key": "SHA256(UTF-8 canonical JSON [seed,source_path,zero_based_row_index]); JSON ensure_ascii=False, separators=(',', ':'); ascending digest bytes",
  "target_tokens": 100000000000
}

Attribution and use

This dataset derives from HuggingFaceFW FineWeb-Edu, released under ODC-By 1.0. See the original dataset card for collection, filtering, limitations, and research attribution, and the Common Crawl terms referenced there. Underlying web content may retain its original authors' rights. This processing changes only the tabular representation and token-count metadata.