CL-From-Nothing/RLVE-Test-Qwen3-1.7B-SFT-warmup-Pass8
RLVE test eval — SFT warmup (pass@8) Evaluation rollouts on the RLVE test split. Model: Qwen3-1.7B-SFT-rlve-20K-1epoch (RL-warmup baseline, pre-RL) Source prompts: RLVE test split — 180 questions (RLVE-Eval Gym environments) Sampling: 8 samples/question (pass@8) = 1440 records, temperature 0.7, max 16384 new tokens Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]). Record-level accuracy (reward>0): 176 / 1440 = 12.2%, mean reward -0.724 pass@8 (>=1 of 8… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Test-Qwen3-1.7B-SFT-warmup-Pass8.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face