YangyiH/qwen3-4b-teacher-rollouts-76k-nonthinking
Qwen3-4B Teacher Rollouts 76K Non-Thinking This dataset contains 76,800 fixed teacher trajectories generated for a prompt-aligned reproduction study of on-policy distillation with Qwen3-1.7B. It is an independent research artifact, not an official release from the model or paper authors. Models and generation Teacher: Qwen/Qwen3-4B-Instruct-2507 Tokenizer/chat template: Qwen/Qwen3-1.7B Mode: non-thinking (enable_thinking=False) Temperature: 0.7 Top-p: 1.0 Top-k:… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/qwen3-4b-teacher-rollouts-76k-nonthinking.
Qwen3-4B Teacher Rollouts 76K Non-Thinking
This dataset contains 76,800 fixed teacher trajectories generated for a prompt-aligned reproduction study of on-policy distillation with Qwen3-1.7B. It is an independent research artifact, not an official release from the model or paper authors.
Models and generation
- Teacher:
Qwen/Qwen3-4B-Instruct-2507 - Tokenizer/chat template:
Qwen/Qwen3-1.7B - Mode: non-thinking (
enable_thinking=False) - Temperature: 0.7
- Top-p: 1.0
- Top-k: 0
- Maximum prompt length: 4,096 tokens
- Maximum completion length: 8,192 tokens
- Seed for row
i:42 + i
Prompts were passed to the vLLM completions API as exact token IDs with no additional special tokens. Returned token IDs and special tokens are retained.
Statistics
The first 25,600 trajectories exactly reuse the earlier 100-step teacher bank. The remaining 51,200 rows extend the same seed-42 prompt permutation without an epoch wrap.
Fields
input_ids: concatenated prompt and completion token IDslabels:-100over prompt tokens and completion token IDs thereafterprompt_ids,completion_ids: separate exact token sequencesmessages: original conversation plus decoded teacher assistant responsedomain: math, science, or codesource_dataset,source_config,source_split,source_shard,source_id: upstream provenanceprompt_sha256: hash of the source promptsource_row_index,unified_order_index: source and rollout schedule indicesfinish_reason,generation_seed: generation metadata
Integrity
- Prompt order SHA-256:
c1cc096fa94f82bf44457673b112039f604aaba19711a850f368514da6d7429d - Completion IDs SHA-256:
a3c5ae8551b23025bb57d6d60a011ef7e156d034d0853656dbb33d29b064127b
The repository includes the full generation manifest. A post-generation audit found zero prompt-identity mismatches, and the prompt-token hash matches a fresh application of the documented Qwen non-thinking chat template.
Usage
from datasets import load_dataset
dataset = load_dataset(
"YangyiH/qwen3-4b-teacher-rollouts-76k-nonthinking",
split="train",
)Source data and licensing
Prompts originate from the NVIDIA OpenMathReasoning, OpenScienceReasoning-2, and OpenCodeReasoning datasets. This repository does not assert a new blanket license over upstream prompt content. Users should review and comply with the source datasets' licenses, terms, and attribution requirements, as well as the Qwen model license.
Limitations
- The source mixture intentionally retains duplicate prompts inherited from the upstream datasets.
- No additional quality, difficulty, benchmark-contamination, or safety filtering was applied.
- 14.29% of completions reached the configured 8,192-token cap and may be truncated rather than naturally terminated.
- Generated responses can contain errors or undesirable content inherited from the model and source prompts.
