datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-1.5-RL-Verifiable
Dataset Card for NuminaMath-1.5-RL-Verifiable
Dataset Summary
NuminaMath-1.5-RL-Verifiable is a curated subset of the NuminaMath-1.5 dataset, specifically filtered to support reinforcement learning applications requiring verifiable outcomes. This collection consists of 131,063 math word problems from the original dataset that meet strict filtering criteria: all problems have definitive numerical answers, validated problem statements and solutions, and come from… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-RL-Verifiable.a11oy-verifiable-corpus
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
a11oy — Verifiable Corpus · verify it yourself
This dataset publishes a11oy's signed receipts and proof surface so that
anyone can independently verify them — no trust in SZL Holdings required.
Every receipt here carries the full cryptographic material needed to check its
signature offline;… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/a11oy-verifiable-corpus.verifiable-coding-problems
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
verifiable-coding-problems-python
Dataset Card for Verifiable Coding Problems Python 10k
This dataset contains all Python problems from PrimeIntellect's verifiable-coding-problems dataset. We have formatted the verification_info and metadata columns to be proper dictionaries, but otherwise the data is the same. Please see their dataset for more details.
verifiable-code-reasoning
Verifiable Code Reasoning
Execution-verified Python problems with chain-of-thought
Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text
Overview
Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests.
Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if:
a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.verifiable-coding-problems-python_decontaminated-testedverifiable-coding-problems-python_decontaminatedverifiable-math-problems
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
verifiable-coding-problems-python_decontaminated-tested-shuffledVerifiableQFT
Verifiable Synethetic QFT Problems
This dataset provides the synthetic QFT problems and rejection sampled CoT samples used in Fine-Tuning Small Reasoning Models for Quantum Field Theory by N. Woodward et al..
The dataset consists of 2,588 synthetic Quantum Field Theory problems with auto-verifiable code solutions and 24,918 rejection-sampled chain-of-thought (CoT) solutions for supervised fine-tuning.
Dataset Summary
This dataset provides two complementary… See the full description on the dataset page: https://huggingface.co/datasets/nswoodward/VerifiableQFT.pdf_science_questions_verifiable_r1_traces__2_24_25
Dataset card for pdf_science_questions_verifiable_r1_traces__2_24_25
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"success": true,
"page_count": 37,
"page_number": 1,
"question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verifiable_r1_traces__2_24_25.NuminaMath-1.5-Verifiable
NuminaMath-1.5-Verifiable
A filtered subset of NuminaMath-1.5, retaining only non-synthetic examples with valid answers.
Filtering Criteria
• Excludes synthetic examples.
• Keeps only entries with non-empty, meaningful answers.
• Removes generic placeholders like “proof” and “notfound.”
Usage
from datasets import load_dataset
dataset = load_dataset("yentinglin/NuminaMath-1.5-Verifiable")
verifiable-corpus
verifiable-corpus
This is the corpus from "Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning".
Code: https://github.com/jonhue/ttc
Introduction
We study how large language models (LLMs) can continually improve at reasoning on their target tasks at test-time. We propose an agent that assembles a task-specific curriculum, called test-time curriculum (TTC-RL), and applies reinforcement learning to continue training the model for its target task.… See the full description on the dataset page: https://huggingface.co/datasets/lasgroup/verifiable-corpus.repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts
Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985)
Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning
with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) —
OpenReview weMYE1B16x,
arXiv 2602.08499.
Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge.
Official code: github.com/lxd99/CBS_public (verl 0.5.x fork).
What CBS is
The paper reframes rollout scheduling in RLVR as… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts.medical-verifiable-dedup
Medical QA Dataset
Overview
This dataset is a collection of multiple Medical QA sources, benchmarks, mock tests, and extracted data from various PDFs. It is intended for research and development in medical question-answering tasks.
We deduplicated based on UUID from text using MD5.
⚠ Important Note: Only the MedQA and MedMCQA datasets have been filtered to exclude their respective test sets. Other sources may still contain test data, so use caution when evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OpenMedical/medical-verifiable-dedup.verifiable-public-statement-corpus-compiler
Toward a Verifiable Public-Statement Corpus Compiler
This repository publishes the first public edition of a process white paper about an attempted general-purpose system for compiling attributable public statements from public audiovisual media.
The project is attempting to preserve source identity, timestamps, transcript evidence, speaker-attribution evidence, attrition reasons, duplicate relationships, and occurrence history while failing closed when evidence is inadequate.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/verifiable-public-statement-corpus-compiler.magpie-reasoning-v1-20k-math-verifiable-step-by-step-rationale-alpaca-formatpre1900-verifiable-physicsmagpie-reasoning-v1-20k-math-verifiable-step-by-step-rationaleverifiable-coding-problems-python-10k_decontaminatednumina_math_ko_verifiable_540k
Dataset Card: OLAIR/numina_math_ko_verifiable_540k
Overview:A paired dataset of math questions (translated into Korean using GPT-4o-mini) and verifiable answers. Intended for RL training (e.g., GRPO) and mathematical reasoning tasks.
Sources:
Questions: Derived from AI-MO/NuminaMath-CoT
Answers: Extracted from flatlander1024/numinamath_verifiable_cleaned
Key Points:
Translation: No-cleansing version; translations may contain errors.
Usage: Suitable for RL and language… See the full description on the dataset page: https://huggingface.co/datasets/OLAIR/numina_math_ko_verifiable_540k.verifiable-pythonic-function-calling-lite
Verifiable Pythonic Function Calling Lite
This dataset is a subset of pythonic function calling dataset that is used for training Pythonic function calling models Dria-Agent-a-3B and Dria-Agent-a-7B.
Dria is a python framework to generate synthetic data on globally connected edge devices with 50+ models. See the network here
Dataset Summary
The dataset includes various examples of function calling scenarios, ranging from simple to complex multi-turn interactions.
It… See the full description on the dataset page: https://huggingface.co/datasets/driaforall/verifiable-pythonic-function-calling-lite.OpenThoughts2-1M_verifiablejson-mode-verifiableverifiable-reasoning-filtered-gpt-41polaris_filtered_nemotron_medium_math_verifiable
Polaris Filtered Nemotron Medium Sympy Verifiable (v2)
This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using math verify-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset.
Dataset Summary
Total Original Samples: 2,424,392
Final Kept Samples: 357,790 (14.8%)
Target Reasoning Length: 4k-8k… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_math_verifiable.mgsm8k-instruct-verifiableverifiable-reasoning-filtered-o4-miniverifiable-coding-problems
