CoolFace
Datasetpublic

bicycleman15/fineweb_50b_none_none

fineweb_50b_none_none Tokenized FineWeb-Edu sample-100BT (no document-length filter). Tokenizer: NousResearch/Llama-2-7b-hf (Llama-2, vocab 32k) Format: uint16 NumPy shards (shard_train_*.npy, shard_val_*.npy) Size: 50B train tokens, 100M validation tokens, 100M tokens per shard Each document is prefixed with the EOS token License: ODC-By 1.0 (same as FineWeb-Edu). Also subject to Common Crawl Terms of Use.

sourceHugging Faceodc-byupdated 22d agoView on Hugging Face
0likes539downloads
Dataset Card

fineweb50bnone_none

Tokenized FineWeb-Edu sample-100BT (no document-length filter).

  • —Tokenizer: `NousResearch/Llama-2-7b-hf` (Llama-2, vocab 32k)
  • —Format: uint16 NumPy shards (shard_train_*.npy, shard_val_*.npy)
  • —Size: 50B train tokens, 100M validation tokens, 100M tokens per shard
  • —Each document is prefixed with the EOS token

License: ODC-By 1.0 (same as FineWeb-Edu). Also subject to Common Crawl Terms of Use.