rlve
OpenThinker3-1.5B-RLVE-GGUFNemotron-Research-Reasoning-Qwen-1.5B-v2-RLVE-i1-GGUFOpenThinker3-1.5B-RLVE-i1-GGUFNemotron-Research-Reasoning-Qwen-1.5B-v2-RLVE-GGUFrose-olmo3-7b-from-qwen3-32b-rlve-k1024-step140Nemotron-Research-Reasoning-Qwen-1.5B-v2-RLVEOpenThinker3-1.5B-RLVEqwen3-30b-a3b-0703-fixed-rlve-official16-iter351
Datasets
All datasets matching “rlve”RLVE-Qwen3-1.7B-Pass1-Rollouts
RLVE teacher rollouts — Qwen3-1.7B (pass@1)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-1.7B
Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym
environments (counting / combinatorics / optimization tasks)
Sampling: 1 sample/question (pass@1) = 9000 records,
temperature 0.7, max 4096 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.RLVE_envs65_data1000rlve_teacherrl-verifier-pitfalls_hacking_dataThis is the hacking dataset associated with the paper "Pitfalls of Rule- and Model-based Verifiers -- A Case Study on Mathematical Reasoning." In this paper, we design and release a "Hacking Dataset" of 13 adversarial patterns (e.g., gibberish, HTML tags, empty symbols) to evaluate verifier robustness.
rlve_recoverablerlve-eval-rollouts
