birgermoell/oellm-longctx-tokenized-natural-128k-256k-pilot-v1
OELLM Natural Long-Context Tokenized 128K/256K Expanded Pilot This dataset contains Megatron-LM tokenized continuation-training artifacts for natural long-context extension experiments at 128K and 256K sequence scales. This public revision expands the original pilot from 128 to 512 packed examples per source/tier, for 3,072 packed examples total. The raw pack manifest reports approximately 586M source-side estimated tokens across all six source/tier shards. Accessible source… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-natural-128k-256k-pilot-v1.
Fix expanded dataset card formatting
Expand natural 128K/256K pilot to 512 examples per source tier
Upload natural 128K/256K pilot Megatron artifacts
initial commit
