CoolFace
Datasetpublic

BoomQ/fineweb-edu-2016-qwen2-sample

FineWeb-Edu 2016 / Qwen2 Consistency sample — not the completed year. Documents: 900. Actual recounted Qwen2 tokens: 937,977. Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9. The input inventory covers 9 crawl directories. date is the integer crawl year 2016, not an article publication date. Original text is preserved without cleaning, normalization, deduplication, truncation, or added formatting. Source token counts are not used. Optional… See the full description on the dataset page: https://huggingface.co/datasets/BoomQ/fineweb-edu-2016-qwen2-sample.

sourceHugging Faceodc-byupdated 15d agoView on Hugging Face
0likes79downloads
Dataset Card

FineWeb-Edu 2016 / Qwen2

Consistency sample — not the completed year.

Documents: 900. Actual recounted Qwen2 tokens: 937,977.

Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9. The input inventory covers 9 crawl directories. date is the integer crawl year 2016, not an article publication date. Original text is preserved without cleaning, normalization, deduplication, truncation, or added formatting. Source token counts are not used. Optional source/category columns are omitted.

Tokenizer: Qwen/Qwen2-7B, revision 453ed1575b739b5b03ce3758b23befdb0967f40e, use_fast=True, trust_remote_code=False; Transformers 5.0.0, tokenizers 0.22.2. Each raw text is tokenized with add_special_tokens=False, truncation=False, padding=False; attention masks and token type IDs are disabled. The saved count is the length of input_ids. Documents beyond the model context length remain intact.

Full provenance, exact counts, file checksums, and selection details are in manifest.json. The pinned input file inventory is in source_manifest.json.

This review sample contains the first 100 documents from the first lexicographic shard in each crawl. It is a crawl-coverage sample, not an unbiased estimate of the year's token total. Every saved sample text was matched to its source and independently recounted after Parquet reload. source_rows.json records the source row and per-document text SHA256 for each sample row.

Attribution and use

This dataset derives from HuggingFaceFW FineWeb-Edu, released under ODC-By 1.0. See the original dataset card for collection, filtering, limitations, and research attribution, and the Common Crawl terms referenced there. Underlying web content may retain its original authors' rights. This processing changes only the tabular representation and token-count metadata.