CL-From-Nothing/RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts
RLVE teacher rollouts — Qwen3-4B-Thinking-2507 (pass@8) Teacher rollouts for on-policy distillation on the RLVE environment suite. Teacher / sampler: Qwen3-4B-Thinking-2507 Source prompts: RLVE train split — 9000 questions across 18 environments (counting / combinatorics / optimization tasks) Sampling: 8 samples/question (pass@8) = 72000 records, temperature 1.0 (sample.sh default 0.7 -> here T per run), max 16384 new tokens Rewards: recomputed offline with the RLVE-Eval Gym… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts.
022
RLVE teacher rollouts — Qwen3-4B-Thinking-2507 (pass@8)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
- Teacher / sampler: Qwen3-4B-Thinking-2507
- Source prompts: RLVE
trainsplit — 9000 questions across 18 environments (counting / combinatorics / optimization tasks) - Sampling: 8 samples/question (
pass@8) = 72000 records, temperature 1.0 (sample.sh default 0.7 -> here T per run), max 16384 new tokens - Rewards: recomputed offline with the RLVE-Eval Gym verifier (
reward in [-1, 1]: +1 correct, -0.5 wrong answer, -1 wrong/missing format). Teacher accuracy (reward>0): 23153 / 72000 = 32.2%.
