datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlvr-code-data-python-r1-format-filteredRLVRAMBench
RLVRAMBench
Which language-model training configurations can I use with the memory
I have, and how much testing does that decision require?
RLVRAMBench is a measurement dataset with open evaluation tasks for a
specific language-model training system. It measures memory feasibility
when response generation and reinforcement-learning updates share the
same graphics processors. It provides measured outcomes, fixed prediction
tasks, a budgeted decision replay, reference methods, and… See the full description on the dataset page: https://huggingface.co/datasets/kobzaond/RLVRAMBench.rlvr-prompts_responses-mixin_it_up-v2-filtered-no-chineserlvr_mixin_it_up_prompts-qwen3-32b-06B-thoughts-x8-filtered-no-chineserlvr-code-data-python-r1rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.oellm-code-rlvr
OpenEuroLLM Code RLVR
oellm-code-rlvr is a deterministic corpus of 100,000 Python programming prompts for reinforcement learning with verifiable rewards. Every task uses standard input/output, includes two model-visible examples, and has 10–13 hidden tests in the Open R1 verification_info format.
The corpus is procedural and Apache-2.0 licensed. It does not copy Codeforces, LeetCode, LiveCodeBench, HumanEval, MBPP, APPS, or other benchmark text.
Design
The release… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-code-rlvr.chem-rlvr-TEST
ChemBench-RLVR: Comprehensive Chemistry Dataset for Reinforcement Learning from Verifiable Rewards
Dataset Description
ChemBench-RLVR is a high-quality, balanced dataset containing 7,001 question-answer pairs across 14 chemistry task types. This dataset is specifically designed for training language models using Reinforcement Learning from Verifiable Rewards (RLVR), where all answers are computationally verifiable using established cheminformatics tools.
Key… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chem-rlvr-TEST.Chem-RLVR
Chem-RLVR
Chem-RLVR is a benchmark for reasoning over experimental reaction records,
with direct reaction-yield prediction and counterfactual yield-reasoning tasks.
Paper: "Chem-RLVR: Verifier-Based Training for Reaction Yield Prediction"
Dataset
Chem-RLVR contains 12,000 questions:
6,000 yield-prediction questions
6,000 counterfactual questions
Each task contains:
4,200 RL-training examples
1,800 frozen held-out evaluation examples
The benchmark covers three… See the full description on the dataset page: https://huggingface.co/datasets/redugo/Chem-RLVR.aime2024-25-rlvr-olmo3-7b-base-pass64-quartilesoellm-math-rlvr
OpenEuroLLM Math RLVR
One million deterministic, verifier-ready mathematical problems for reinforcement learning with
verifiable rewards. The release contains a 760,000-row English depth pool and 10,000 aligned semantic
problems rendered in all 24 official EU languages (240,000 rows).
This is a prompt-and-answer rollout corpus, not a chain-of-thought corpus. Model inputs contain only the
problem and output-format instruction. Reference answers and verifier contracts remain… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-math-rlvr.rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source
Tool-Reasoning SFT — RLVR Retrieval Source Trajectories
156,381 multi-turn agentic retrieval trajectories across three document corpora, in a strict reasoning + tool-call format with validated FSM transitions. Each trajectory records a model searching a corpus, opening documents, and citing relevant passages to answer a question.
Author: Aman Priyanshu
Source Environments
Trajectories were collected against three RLVR retrieval environments from the FORMAT: Search -… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source.countdown-rlvr
Countdown RLVR
Qwen3-4B의 검증 가능한 추론 학습에 사용하는 Countdown 데이터셋입니다.
주어진 숫자를 각각 한 번 사용하여 목표값을 만드는 수식을 생성합니다.
1. 데이터 구성
분할
개수
숫자 개수
목표값
SHA-256
train
1,024
4
10~100
aa7abb6242d8ada72e55a6d8d0917e3618473ddf2f0f880b288814394b231131
validation
128
4
10~100
b06a1be3604d637aa19bd61af57aadf98fbbffcb8ef4db8d477fe8535c617497
test
256
4
10~100
416c02076321875cccfeed19f742e56048269b4b9d24112f6a2bee82ba301815
demo.jsonl에는 검증 흐름을 확인하는 숫자 3개 문제를 둡니다.… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/countdown-rlvr.turkish-math-rlvr
Turkish Math Reasoning Dataset for RLVR training
What this dataset is
This dataset is a Turkish math reasoning benchmark augmented with a weak-model pass-rate difficulty signal, designed for curriculum learning, GRPO / RLVR-style training, and evaluation of reasoning models in Turkish math problems.
It is constructed by merging two Turkish math datasets and annotating each problem with a pass rate computed by the google/gemma-3-4b-it model.
Each example represents one… See the full description on the dataset page: https://huggingface.co/datasets/barandinho/turkish-math-rlvr.rlvr-code-data-python-r1-format-filteredrlvr-bash-terminal-bench
rlvr-bash-terminal-bench
RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks.
Stats
Metric
Value
Total samples
1,120
Unique tasks
88
Avg samples/task
12.7
Average reward
0.249
Perfect solutions (reward=1.0)
10.4%
Partial solutions (0<reward<1)
28.8%
Zero reward
60.8%
Tasks fully solved
13.6%
Format
{
"task_id": "string",
"prompt": "string",
"completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.math_rlvr_mixture_dporlvr-code-view-tool-new-first-turn-only-user-with-repo-namecode_rlvr_mixture_dpoturkish-math-rlvr
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti barandinho tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: barandinho/turkish-math-rlvr
🔗 Derleyen Platform: VeriPazarı
Türkçe Matematiksel Akıl Yürütme (RLVR Eğitim Veri Seti)
Bu Veri Seti Nedir?
Bu veri seti, zayıf bir modelin başarı oranına… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-math-rlvr.code_rlvr_mixture_dporlvr_task116_com2sense_commonsense_reasoningfinancial-rlvr-verified-v1
Financial-RLVR-Verified-v1 (Alpha Preview)
Overview
This dataset provides 100% Execution-Verified Reasoning Trajectories for Financial Quant, Corporate Valuation, and Fixed Income analysis. Designed specifically for RLVR (Reinforcement Learning with Verifiable Rewards) algorithms.
Key Features
Deterministic Ground Truth: All numerical outputs are verified inside a sandboxed Python execution environment.
Zero-Hallucination Guarantee: Floating-point… See the full description on the dataset page: https://huggingface.co/datasets/coslinedev/financial-rlvr-verified-v1.qwen3_0.6b-rlvr_task1385_anli_r1_entailmentRLVR-datasets
RLVR Datasets
Dataset files for the VO/RLVR experiment configured in vo/tmp/run_vo.sh.
Files
train/math__combined_54.4k.parquet: training set used by TRAIN_FILE; 54,404 rows.
validation/math__aime_repeated_32x_960.parquet: AIME validation set used by VAL_AIME_FILE; 960 rows.
validation/math__math_500.parquet: MATH-500 validation set used by VAL_MATH500_FILE; 500 rows.
The parquet files preserve the local schemas used by the verl/RLVR training script, including prompt… See the full description on the dataset page: https://huggingface.co/datasets/back-prop/RLVR-datasets.rlvr_task1209_atomic_classification_objectuserlvr_task1211_atomic_classification_hassubeventqwen3_0.6b-rlvr_task073_commonsenseqa_answer_generationqwen3_0.6b-rlvr_task1419_mathqa_gain
