ZhuofengLi/fineweb-edu-pretokenized-10b
FineWeb-Edu Pretokenized 10B Megatron indexed-dataset (.bin/.idx) versions of HuggingFaceFW/fineweb-edu sample/10BT. Each Hugging Face subset contains a small metadata.parquet index. The actual Megatron files are under <subset>/files/; pass each prefix without the .bin/.idx suffix to Megatron Core or Megatron Bridge. Documents retain upstream shard and row order. Tokenization disables automatic special-token insertion and appends exactly one tokenizer EOS/EOD token to each… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-10b.
FineWeb-Edu Pretokenized 10B
Megatron indexed-dataset (.bin/.idx) versions of HuggingFaceFW/fineweb-edu sample/10BT.
Each Hugging Face subset contains a small metadata.parquet index. The actual Megatron files are under <subset>/files/; pass each prefix without the .bin/.idx suffix to Megatron Core or Megatron Bridge.
Documents retain upstream shard and row order. Tokenization disables automatic special-token insertion and appends exactly one tokenizer EOS/EOD token to each document. The data is not shuffled, packed, truncated, or sentence-split.
