datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tulu-v2-sft-mixture-olmo-4096
Dataset Card for Tulu V2 Mix (4096 OLMo version)
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This is a modified version of the Tulu V2 Mix used to train newer (after April 2024) OLMo-SFT/Instruct variants (e.g. this model, or this one).
The only difference is that the hardcoded subset (dataset='hard_coded') has been replaced… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-4096.polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps
Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507
Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's
actual continuation for each, and the token accounting behind it.
The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is
decoded to text, the teacher is shown it under its own chat template, and the teacher's
reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.per-context-rb-l0-4096-no-eos-qwen3-1.7b-compression-bs32-32k-146102-rollouts
per_context_rb_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_1epoch rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
deepseek_grpo_correct_4096
DeepSeek GRPO Correct 4096
Filtered GRPO training subset generated from deepseek-reasoner math generations.
Files
train.jsonl: filtered training examples with prompt, solution, dataset_index, and DeepSeek metadata.
metadata.json: filtering metadata.
Filter
Rows are kept when the raw generation is successful, stopped, correct, deduplicated by dataset_index, and has usage_total_tokens <= 4096.
Summary
Rows: 5502
Max total tokens: 4096
Source raw… See the full description on the dataset page: https://huggingface.co/datasets/igreck/deepseek_grpo_correct_4096.minesweeper-student-minekuk-qwen1.7b-continued-by-qwen3-4b-thinking-t4096-r16384oryzaModel4096minesweeper-Qwen_Qwen3-4B-Thinking-continued-by-Qwen_Qwen3-1.7B-trunc4096-resp16384minesweeper-student-kukurasu20k-qwen1.7b-e3-mask-continued-by-qwen3-4b-thinking-t4096-r16384webshop_success_len_lt4096minesweeper-qwen3-4b-thinking-continued-by-teacher-kukurasu20k-qwen1.7b-e3-mask-t4096-r16384sudoku-Qwen3-4BThinking-contby-Qwen_Qwen3-1.7B-trunc4096-resp16384kukurasu-student-minekuk-qwen1.7b-continued-by-qwen3-4b-thinking-t4096-r16384sudoku-Qwen3-4BThinking-contby-tp-mswp-kuku-nemotron-cascade-8b-trunc4096-resp16384minesweeper-Qwen_Qwen3-4B-Thinking-continued-by-nvidia_Nemotron-Cascade-8B-trunc4096-resp16384ruozhiba_zhminesweeper-allenai_OLMo-3-7B-Think-continued-by-Qwen_Qwen3-4B-trunc4096minesweeper-qwen3-4b-thinking-continued-by-teacher-kukurasu20k-nemotron-e3-mask-t4096-r16384minesweeper-Qwen_Qwen3-1.7B-continued-by-Qwen_Qwen3-4B-trunc4096minesweeper-Qwen_Qwen3-1.7B-continued-by-Qwen_Qwen3-4B-trunc4096-resp16384minesweeper-allenai_OLMo-3-7B-Think-continued-by-Qwen_Qwen3-4B-Thinking-trunc4096-resp16384minesweeper-student-minekuk-nemtron8b-continued-by-qwen3-4b-thinking-t4096-r16384riceCad-4096-14speciesminesweeper-Qwen_Qwen3-1.7B-continued-by-Qwen_Qwen3-4B-Thinking-trunc4096-resp16384kukurasu-Qwen_Qwen3-4B-Thinking-continued-by-Qwen_Qwen3-1.7B-trunc4096-resp16384sudoku-nvidia_Nemotron-Cascade-8B-continued-by-Qwen_Qwen3-4B-Thinking-trunc4096-resp16384rlve-multitask-qwen3-4b-n4-randcut512-4096x20-completed-by-qwen3-4b-thinking-r16384minesweeper-nvidia_Nemotron-Cascade-8B-continued-by-Qwen_Qwen3-4B-trunc4096kukurasu-Qwen_Qwen3-1.7B-continued-by-Qwen_Qwen3-4B-Thinking-trunc4096-resp16384kukurasu-Qwen_Qwen3-4B-Thinking-continued-by-nvidia_Nemotron-Cascade-8B-trunc4096-resp16384sudoku-Qwen_Qwen3-1.7B-continued-by-Qwen_Qwen3-4B-Thinking-trunc4096-resp16384
