datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.RLVR-IFeval
IF Data - RLVR Formatted
This dataset contains instruction following data formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards.
Prompts with verifiable constraints generated by sampling from the Tulu 2 SFT mixture and randomly adding constraints from IFEval.
Part of the Tulu 3 release, for which you can see models here and datasets here.
Dataset Structure
Each example in the dataset contains the standard instruction-tuning… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-IFeval.Multi-subject-RLVRMulti-subject data for paper "Expanding RL with Verifiable Rewards Across Diverse Domains".
we use a multi-subject multiple-choice QA dataset ExamQA (Yu et al., 2021).
Originally written in Chinese, ExamQA covers at least 48 first-level subjects.
We remove the distractors and convert each instance into a free-form QA pair.
This dataset consists of 638k college-level instances, with both questions and objective answers written by domain experts for examination purposes.
We also use GPT-4o-mini… See the full description on the dataset page: https://huggingface.co/datasets/virtuoussy/Multi-subject-RLVR.math-rlvr-mini-smollm2-0.4b-v2rlvr-reward-hacking-scale-no-conftest-20260909
Matched no-conftest RLVR study 20260909
Complete immutable training, monitoring and comparison trajectories for six models.
All valid outcomes are retained, including refusals, failures and truncations.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training, monitoring, comparison, or a derived judgment.
import json
from datasets import load_dataset
rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.open-reasoning-rlvr-24k
Open Reasoning RLVR mixture (math : science : code = 1 : 1 : 1)
A verifier-carrying 1:1:1 subsample of NVIDIA's open reasoning corpora,
built for Dr.GRPO / RLVR runs — 8000 train and 500 validation prompts
per domain.
domain
source
verifier
signal
math
nvidia/OpenMathReasoning (cot)
math_boxed
\boxed{} vs answer
science
nvidia/OpenScienceReasoning-2
mcq_boxed
\boxed{} option letter vs answer
code
nvidia/OpenCodeReasoning (split_0)
stdio_tests
program run on the… See the full description on the dataset page: https://huggingface.co/datasets/Ksgk-fy/open-reasoning-rlvr-24k.RLVR-GSM-MATH-IF-Mixed-Constraints
GSM/MATH/IF Data - RLVR Formatted
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data.
This dataset contains data formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards.
It was used to train the final Tulu 3 models with RL, and contains the following subsets:
GSM8k (7,473 samples): The GSM8k train set formatted for use with RLVR and open-instruct. MIT License.
MATH (7,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-GSM-MATH-IF-Mixed-Constraints.hle_rlvr_no_promptrlvr-code-data-python-r1-format-filteredGPQA-train-RLVRrlvr-guru-raw-data-extended
RLVR GURU Extended: Compiling a 150K Cross-Domain Dataset for RLVR
A comprehensive cross-domain reasoning dataset containing 150,000 training samples and 221,332 test samples across diverse reasoning-intensive domains. This dataset extends the foundational work from the GURU dataset (Cheng et al., 2025) by incorporating additional STEM reasoning domains (MedMCQA and CommonsenseQA) while maintaining rigorous quality standards and verification mechanisms essential for reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/rlvr-guru-raw-data-extended.math-rlvr-mini
Math RLVR Mini
227,443 deduplicated, numeric-verifiable math problems from
7 source lineages. Rows expose exactly
{source, question, ground_truth} and preserve parent order.
Decontamination
Starting from cs-giung/math-rlvr-mini@f6540bd1e1118932b9c32f7e282d397788334ded,
this revision is checked against all 19 splits
and 8,774 rows in cs-giung/math-evals
revision 07268cc7bfd36cce0d390ec337b8377dbced4b94, plus cs-giung/math-think-sft-mini
revision… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/math-rlvr-mini.RLVRAMBench
RLVRAMBench
Which language-model training configurations can I use with the memory
I have, and how much testing does that decision require?
RLVRAMBench is a measurement dataset with open evaluation tasks for a
specific language-model training system. It measures memory feasibility
when response generation and reinforcement-learning updates share the
same graphics processors. It provides measured outcomes, fixed prediction
tasks, a budgeted decision replay, reference methods, and… See the full description on the dataset page: https://huggingface.co/datasets/kobzaond/RLVRAMBench.RLVR-GSM
GSM8k Data - RLVR Formatted
This dataset contains the GSM8k dataset formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards.
Part of the Tulu 3 release, for which you can see models here and datasets here.
Dataset Structure
Each example in the dataset contains the standard instruction-tuning data points as follow:
messages (list): inputs used to prompt the model (after chat template formatting).
ground_truth (str): the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-GSM.rlvr-prompts_responses-mixin_it_up-v2-filtered-no-chineserlvr_function_type_inferencerlvr_gsm8k_zsRLVR-MATHnemotron-rlvr-openinstructrlvr_mixin_it_up_prompts-qwen3-32b-06B-thoughts-x8-filtered-no-chineserlvr-code-data-python-r1GPQA-RLVROne-Shot-RLVR-DatasetsThis repository contains the dataset presented in Reinforcement Learning for Reasoning in Large Language Models with One Training Example.
Code: https://github.com/ypwang61/One-Shot-RLVR
chem-rlvr-TEST-4
ChemBench-RLVR: Comprehensive Chemistry Dataset for Reinforcement Learning from Verifiable Rewards
Dataset Description
ChemBench-RLVR is a high-quality, balanced dataset containing 16,699 question-answer pairs across 14 chemistry task types. This dataset is specifically designed for training language models using Reinforcement Learning from Verifiable Rewards (RLVR), where all answers are computationally verifiable using established cheminformatics tools.
Key… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chem-rlvr-TEST-4.math-rlvr-mini-sa-smollm2-0.1b-v1rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.math-rlvr-mini-sa-smollm2-0.1b-v0cap-rlvr-entail
CAP RLVR Case Relationship Classification Dataset
Determining how cases relate (overrule, distinguish, affirm, etc.)
Dataset Overview
Task Type: Case Relationship Classification
Train samples: 5,159,319
Validation samples: 644,914
Test samples: 644,914
Estimated size: 5GB
Task Description
Determining how cases relate (overrule, distinguish, affirm, etc.)
Data Format
Each record contains:
inputs: The question or prompt for the legal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/kylebrussell/cap-rlvr-entail.Qwen3-0.6B-cybertown-RLVR-data
Qwen3-0.6B-cybertown-RLVR-data
This dataset contains Cybertown RLVR training and validation data used for WindyLab/Qwen3-0.6B-cybertown-RLVR.
Files
train.parquet: RLVR training split.
val.parquet: RLVR validation split.
manifest.json: dataset construction metadata.
states.index.json: index mapping state ids to sharded state records.
states.shards.manifest.json: state shard metadata.
states/: sharded replan state records.
The legacy monolithic states.jsonl is… See the full description on the dataset page: https://huggingface.co/datasets/WindyLab/Qwen3-0.6B-cybertown-RLVR-data.math-rlvr-mini-sa-smollm2-0.1b-v2
