GaloisTheory123/msm-v2-shared-c4-36k
MSM v2 shared C4 36k Training-ready inputs for the second MSM run. Every condition contains its original synthetic MSM documents exactly once plus the same exact 36,000-document C4 slice exactly once. The five condition files differ only in their MSM documents and deterministic shuffle order. Synthetic rows begin with <DOCTAG>\n and declare the same string in mask_prefix; C4 rows are untagged and declare an empty mask_prefix. The trainer must mask only the declared prefix tokens… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/msm-v2-shared-c4-36k.
MSM v2 shared C4 36k
Training-ready inputs for the second MSM run. Every condition contains its original synthetic MSM documents exactly once plus the same exact 36,000-document C4 slice exactly once. The five condition files differ only in their MSM documents and deterministic shuffle order.
Synthetic rows begin with <DOCTAG>\n and declare the same string in mask_prefix; C4 rows are untagged and declare an empty mask_prefix. The trainer must mask only the declared prefix tokens from the language-modeling loss. All document content and EOS tokens remain supervised.
See training_manifest.json for immutable upstream revisions, tokenizer-based token counts, shuffle seeds, file hashes, and realized C4 token fractions. The standalone shared corpus is shared/c4_en_36k.jsonl.gz; directly trainable files are under training/.
