Verification
distill_r1_qwen_math_1.5b_128_solns_math_verificationsatlas-31-strengthening-candidate-verification-under-rl
31. Strengthening candidate verification under reinforcement learning
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 31-strengthening-candidate-verification-under-rl/; the saved training steps and the per-token training arrays are on the Hugging Face repository t2ance/atlas-31-strengthening-candidate-verification-under-rl only.
How can reinforcement learning make the orchestrator's comparing and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-31-strengthening-candidate-verification-under-rl.distill_qwen_7b_aime_verifications_7b_ft_verifierbitaudit_verification_dataset_v2vistr-process-verification-pilot
ViSTR Process-Verification Pilot (14 answer-correct trajectories, multimodal)
Agent trajectories for studying process false positives in multimodal agents:
cases where the answer is correct but the visual reasoning that produced it is
wrong. Ships the raw perception tool outputs so any claim in a trajectory can be
independently re-verified, plus human annotations and an unmodified XSkill
critique of the same trajectories.
Why this exists
Harness / skill… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-process-verification-pilot.HLE-Verifications
HLE with Gemini 3 Pro
This dataset contains 649 multiple-choice and exact-match questions from the Humanity's Last Exam (HLE) benchmark with 50 candidate responses generated by Gemini 3 Pro for each problem. Each response has been evaluated for correctness using a mixture of Qwen3-Next-80B-A3B-instruct and Python code to parse different answer formats, and scored by multiple LLM judges according to a 0-5 rubric.
Dataset Structure
Split: Single split named "data"
Number… See the full description on the dataset page: https://huggingface.co/datasets/FUSE-verifiers/HLE-Verifications.
