gdn2
Datasets
All datasets matching “gdn2”gdn2-ruler-niah-eval-data
RULER NIAH eval data (GDN-2 CPT comparison)
Exact test sets generated with lm-eval-harness RULER generators (RANDOM_SEED=42, tokenizer TinyLlama/TinyLlama_v1.1, lengths [1024, 2048, 4096, 8192], 500 samples/length/task).
Tasks: niah_single_1, niah_single_2, niah_single_3, niah_multikey_1
Used by the unified evaluation in dsc/mc_sketch_remoe/scripts/run_eval_compare_lmeval.sh (limit 50, seed 42). Text-only (raw prompts/targets); each model tokenizes with its own tokenizer.
gdn2-cpt-token-cachegdn2-cpt-fineweb-edu-30k
gdn2-370m CPT pack — gdn2-cpt-fineweb-edu-30k
Pretraining-corpus pack used for the GDN-2 370M SSKetch+ReMoE 4-way CPT runs (2026-08-23).
Source & attribution
Origin dataset: HuggingFaceFW/fineweb-edu config='sample-100BT'
(public; used under the terms stated on its dataset card)
Sampling recipe: streaming skip=800,000, num_samples=30,000 documents
Tokenizer: TinyLlama/TinyLlama_v1.1
Total: 33,569,012 tokens across 30,000 documents
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gyung/gdn2-cpt-fineweb-edu-30k.gdn2-cpt-longdata-30k
gdn2-370m CPT pack — gdn2-cpt-longdata-30k
Pretraining-corpus pack used for the GDN-2 370M SSKetch+ReMoE 4-way CPT runs (2026-08-23).
Source & attribution
Origin dataset: emozilla/Long-Data-Collections-Pretrain-Without-Books
(public; used under the terms stated on its dataset card)
Sampling recipe: streaming skip=800,000, num_samples=30,000 documents
Tokenizer: TinyLlama/TinyLlama_v1.1
Total: 281,330,658 tokens across 30,000 documents
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gyung/gdn2-cpt-longdata-30k.
