CoolFace
Datasetpublic

JackHsieh/dclm-replay.seq-4096.n-262144-olmo3

dclm-replay.seq-4096.n-262144-olmo3 OLMo 3-tokenized CPT-replay sequences for prestar, the OLMo analogue of JackHsieh/dclm-replay.seq-4096.tokens-32B (Qwen3). Source: mlfoundations/dclm-baseline-1.0, pin global-shard_01_of_10/local-shard_0_of_10/*.jsonl.zst. Tokenizer: allenai/Olmo-3-1025-7B; EOD token id 100257 (<|endoftext|>). 262,144 sequences of exactly 4096 tokens each (docs concatenated and packed; EOD-separated). Same builder/pin as the Qwen3 replay — corpus is the same… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/dclm-replay.seq-4096.n-262144-olmo3.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes70downloads
Dataset Card

dclm-replay.seq-4096.n-262144-olmo3

OLMo 3-tokenized CPT-replay sequences for prestar, the OLMo analogue of JackHsieh/dclm-replay.seq-4096.tokens-32B (Qwen3).

  • —Source: mlfoundations/dclm-baseline-1.0, pin global-shard_01_of_10/local-shard_0_of_10/*.jsonl.zst.
  • —Tokenizer: allenai/Olmo-3-1025-7B; EOD token id 100257 (<|endoftext|>).
  • —262,144 sequences of exactly 4096 tokens each (docs concatenated and packed; EOD-separated).
  • —Same builder/pin as the Qwen3 replay — corpus is the same, packing is OLMo-native (sequences do not align 1:1 with the Qwen3 replay).