CoolFace
Datasetpublic

procmarco/fpabl1-arm-a-web-tokens-48k

fpabl1-arm-a-web-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-a-web. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fw2_ita: 2,000,000,000… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-a-web-tokens-48k.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes10downloads
Dataset Card

fpabl1-arm-a-web-tokens-48k

Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-a-web.

  • Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121).
  • Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total.
  • Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks

Source token composition:

  • fw2_ita: 2,000,000,000
  • fwedu: 2,000,000,000
  • replay: 1,000,000,000

See meta.json for the full manifest. Companion arm: the other fpabl1-arm-*-tokens-48k repo.