datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olmo-3-preference-mix-deltas_reasoning-yolo_scottmix-DECON-multi-turnWebShop-Olmo3-7B-Think-Adaptive-Pivot-8192-evalWebShop-Olmo3-7B-Adaptive-Midpoint-eval
WebShop OLMo-3-7B-Instruct adaptive MIDPOINT: held-out evaluation
Held-out WebShop evaluation rows for the checkpoints in
wckwan/WebShop-Olmo3-7B-Adaptive-Midpoint
(merged_hf_actor_gs<N>, appended as the still-training run publishes checkpoints).
Layout
global_step_N/result.json # metrics, diagnostics, goal set, config, fingerprint, runtime
global_step_N/trajectories.jsonl # per-episode trajectories
global_step_N/eval_seed0.log # eval log… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/WebShop-Olmo3-7B-Adaptive-Midpoint-eval.WebShop-Olmo3-7B-SEED-eval
WebShop OLMo-3-7B-Instruct SEED: held-out evaluation
Held-out WebShop evaluation rows for the checkpoints in
wckwan/WebShop-Olmo3-7B-SEED
(merged_hf_actor_gs20 through gs400, every 20 steps).
Layout
global_step_N/result.json # metrics, diagnostics, goal set, config, fingerprint, runtime
global_step_N/trajectories.jsonl # per-episode trajectories
global_step_N/eval_seed0.log # eval log
MANIFEST.sha256 # sha256sum -c compatible… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/WebShop-Olmo3-7B-SEED-eval.WebShop-Olmo3-7B-Adaptive-Pivot-evalWebShop-Olmo3-7B-Adaptive-Pivot-Fenced-eval
WebShop OLMo-3-7B-Instruct pivot FENCED: held-out evaluation
WARNING: global_step_140 is the odd row out. It was produced on the NEW action parser (verl-agent commit 451b4969, which reads the action after the final </think>). It is NOT comparable to global_step_20–global_step_120, which ran on the frozen OLD-parser surface 7173787ad6e171c6 (first <action> pair anywhere). config_fingerprint does not distinguish parser surfaces, so the rows look comparable but are not. The old… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/WebShop-Olmo3-7B-Adaptive-Pivot-Fenced-eval.Alfworld-Olmo3-7B-SEED-eval
ALFWorld OLMo-3-7B-Instruct SEED: held-out evaluation
Held-out ALFWorld evaluation rows for wckwan/Alfworld-Olmo3-7B-SEED (merged_hf_actor_gs20 through gs260).
Layout
global_step_N/result.json # metrics, diagnostics, task list, config, fingerprint, runtime
global_step_N/trajectories.jsonl # per-episode trajectories
global_step_N/eval_seed0.log # eval log (not present for every row)
MANIFEST.sha256 # sha256sum -c compatible… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/Alfworld-Olmo3-7B-SEED-eval.stage1_remix_olmo3_books_260Brollouts-olmo32b-rl
rollouts-olmo32b-rl — evaluation rollouts
Model: GRPO ck300 merged from allenai/Olmo-3-1125-32B (adapters: ReasoningRegisters/olmo32b). Tokenizer used for answer positions: allenai/Olmo-3-1125-32B.
Protocol: 32 rollouts per problem (two seeded batches of 16), temperature 0.6, top-p 0.95,
budget 31,744 generated tokens, seed 20260819. Prompts and grader: the paper's repository
(sophicle/reason). Rollout jsonl files are kept as written (one graded rollout per line, with the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo32b-rl.Alfworld-Olmo3-7B-Adaptive-Midpoint-eval
ALFWorld OLMo-3-7B-Instruct adaptive MIDPOINT: held-out evaluation
Held-out ALFWorld evaluation rows for wckwan/Alfworld-Olmo3-7B-Adaptive-Midpoint (merged_hf_actor_gs<N>), appended as checkpoints are evaluated.
Layout
global_step_N/result.json # metrics, diagnostics, task list, config, fingerprint, runtime
global_step_N/trajectories.jsonl # per-episode trajectories
MANIFEST.sha256 # sha256sum -c compatible
Protocol… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/Alfworld-Olmo3-7B-Adaptive-Midpoint-eval.WebShop-Olmo3-7B-GiGPO-Think-8192-evalAlfworld-Olmo3-7B-Adaptive-Pivot-Fenced-eval
ALFWorld OLMo-3-7B-Instruct adaptive-pivot FENCED: held-out evaluation
Held-out ALFWorld evaluation rows for wckwan/Alfworld-Olmo3-7B-Adaptive-Pivot-Fenced (merged_hf_actor_gs20 through gs200, every 20 steps; the run finished at 200/200).
Layout
global_step_N/result.json # metrics, diagnostics, task list, config, fingerprint, runtime
global_step_N/trajectories.jsonl # per-episode trajectories
global_step_N/eval_seed0.log # eval log
MANIFEST.sha256… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/Alfworld-Olmo3-7B-Adaptive-Pivot-Fenced-eval.Webshop-Olmo3-7B-GRPO-Think-8192-evalLACUNA-data-OLMo3-7B-seed42WebShop-Olmo3-7B-Adaptive-Pivot-Fenced-eval-newparser
WebShop OLMo-3-7B-Instruct pivot FENCED: held-out evaluation, NEW action parser
Held-out WebShop evaluation rows for wckwan/WebShop-Olmo3-7B-Adaptive-Pivot-Fenced, steps gs140 through gs200.
NOT comparable to the gs20–gs120 rows in wckwan/WebShop-Olmo3-7B-Adaptive-Pivot-Fenced-eval. Those ran on the frozen OLD-parser surface 7173787ad6e171c6, which executes the first <action> pair anywhere in the response, including actions rehearsed and rejected inside the reasoning. These… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/WebShop-Olmo3-7B-Adaptive-Pivot-Fenced-eval-newparser.d24-midtrain-olmo3-10b-wholedoc
d24 Midtrain — OLMo-3 Dolmino (10B, whole-doc)
A 9.3B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by
taking a fraction of each component of
allenai/dolma3_dolmino_mix-100B-1025.
No length filter and no chunking — every document is kept whole (some are very long: tens of
thousands of tokens). Each component reaches its target, so the realized mix matches OLMo-3's true
proportions (the OLMo-3 target % column ≈ the kept share). For training, the Megatron… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-10b-wholedoc.d24-midtrain-olmo3-5b
d24 Midtrain — OLMo-3 Dolmino (5B, chunked)
A 5.00B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by
taking a fraction of each component of
allenai/dolma3_dolmino_mix-100B-1025.
Unlike the smaller d24-midtrain-olmo3
(which dropped documents over 2048 tokens — silently zeroing the long reasoning-trace and PDF
components), this build chunks long documents into 2048-token windows (decoded back to text),
so every component keeps ~100% of its tokens and reaches… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b.reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.Alfworld-Olmo3-7B-Adaptive-Pivot-eval
ALFWorld OLMo-3-7B-Instruct adaptive-pivot TERSE: held-out evaluation
Held-out ALFWorld evaluation rows for wckwan/Alfworld-Olmo3-7B-Adaptive-Pivot, the terse-selector control for the ALFWorld OLMo-3 pivot cell. That training run crashed at step 77, so only gs20, gs40 and gs60 exist.
Pairs with the FENCED arm in wckwan/Alfworld-Olmo3-7B-Adaptive-Pivot-Fenced-eval, evaluated on the same box, tree and protocol, for a terse-vs-fenced contrast at gs20-60.
Layout… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/Alfworld-Olmo3-7B-Adaptive-Pivot-eval.dolma3-olmo3-corpus-manifest
dolma3-olmo3-corpus-manifest
Unified per-document manifest variant built against the OLMo3 sidecar schema (1.1B docs, 32-column PyArrow schema with topic + format + quality + token count + source shard).
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_olmo3_corpus_manifest… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-olmo3-corpus-manifest.WebShop-Olmo3-7B-GiGPO-eval-ucl-oldparserdolma3_mix-150B-1025-merged-olmo3
dolma3_mix-150B-1025-merged (OLMo3)
Tokenized copy of the Dolma3 150B mix (dolma3_mix-150B-1025-merged) using the OLMo3 / Qwen3.5-base (q35base) tokenizer.
Format
Megatron-LM indexed binaries: paired *_text_document.bin and *_text_document.idx files (448 files total, ~613 GiB).
These are not Hugging Face datasets Arrow/Parquet shards. Load them with Megatron / Megatron-LM indexed dataset readers.
Source name
Local / blob name:… See the full description on the dataset page: https://huggingface.co/datasets/yangwang92/dolma3_mix-150B-1025-merged-olmo3.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.d24-midtrain-olmo3-5b-wholedoc
d24 Midtrain — OLMo-3 Dolmino (5B, whole-doc)
A 4.71B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix at its exact
component proportions, built by taking a fraction of each component of
allenai/dolma3_dolmino_mix-100B-1025.
Documents are kept whole — no length filter, no chunking.
Why whole-doc: for packed pretrain/midtrain the trainer (e.g. Megatron preprocess_data --append-eod
GPTDataset) already concatenates documents and slices them into context-length windows… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b-wholedoc.mmlu_olmo3_contaminationd24-midtrain-olmo3
d24 Midtrain — OLMo-3 Dolmino subsample
A length-filtered (≤2048 GPT-2 tokens) subsample of OLMo 3's Dolma-3 Dolmino mid-train mix,
used for the d24 v1base-olmo3 midtrain. Built by taking a uniform fraction of each
component of allenai/dolma3_dolmino_mix-100B-1025
(so the mix proportions are preserved), then dropping documents >2048 tokens — which naturally
shrinks the long-doc components (reasoning traces, olmOCR PDFs) that don't fit the 2048 context.
5,902,548 documents… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3.WebShop-Olmo3-7B-GRPO-eval-2048
WebShop OLMo-3-7B-Instruct GRPO: held-out evaluation at response 2048
Held-out WebShop evaluation rows for wckwan/Webshop-Olmo3-7B-Outcome-2048 (merged_hf_actor_gs20 through gs400, every 20 steps).
These rows replace the earlier GRPO rows for this cell (fingerprint 9c807b420392). Those were evaluated at max_response_length=512, a quarter of the trained cap of 2048, so the policy's generations were truncated. Do not mix the two series.
Layout… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/WebShop-Olmo3-7B-GRPO-eval-2048.llm101-olmo3-zh-demo-data
llm001 OLMo3-190M-zh Demo Data
为零基础 AI 大模型研发训练营(llm001)L04 提供的学员数据包。
文件说明
路径
大小
说明
tokenizer/
~4MB
48k BPE 中文分词器,全链路用同一个
nano/tokenized_nano.bin
~1GB
预 tokenize 好的 0.5B tokens(uint16, seed=42 从 3.4B 随机采样 ~15%),直接训
tokenized.bin
~6.8GB
完整 tokenize 数据(3.4B tokens, uint16),有资源可自训完整模型
continue/tokenized_continue.bin
~3.5GB
持续预训练的数据
快速用法(Nano 训练)
from huggingface_hub import snapshot_download
path =… See the full description on the dataset page: https://huggingface.co/datasets/cmz1024/llm101-olmo3-zh-demo-data.FineWeb2-100M-olmo3-7b-toksolmo-3-preference-mix-deltas_reasoning-chosen_qwen8b-yolo_scottmix-DECON
