ncylich/temporal-moe-corpus
temporal-moe-corpus Tokenized training corpus for the Temporal-MoE experiments, 31.3 GiB. This repository holds the tokenized corpus in Megatron indexed-dataset format, plus the 16k tokenizer. It is what the training runs actually read. The raw web-crawl text is not hosted here, because it is a byte-reproducible subset of a public dataset. The exact recipe and the checksums needed to verify a reproduction are below. Contents dclm_tokenized/ 22 x… See the full description on the dataset page: https://huggingface.co/datasets/ncylich/temporal-moe-corpus.
This repository belongs to ncylich on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
