CoolFace
Datasetpublic

birgermoell/oellm-longctx-tokenized-natural-128k-256k-pilot-v1

OELLM Natural Long-Context Tokenized 128K/256K Expanded Pilot This dataset contains Megatron-LM tokenized continuation-training artifacts for natural long-context extension experiments at 128K and 256K sequence scales. This public revision expands the original pilot from 128 to 512 packed examples per source/tier, for 3,072 packed examples total. The raw pack manifest reports approximately 586M source-side estimated tokens across all six source/tier shards. Accessible source… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-natural-128k-256k-pilot-v1.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes30downloads
4 commits on main
49960ff3mo ago

Fix expanded dataset card formatting

birgermoell
4cc61013mo ago

Expand natural 128K/256K pilot to 512 examples per source tier

birgermoell
0e9d09b3mo ago

Upload natural 128K/256K pilot Megatron artifacts

birgermoell
acc20973mo ago

initial commit

birgermoell