datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-1.5-RL-Verifiable
Dataset Card for NuminaMath-1.5-RL-Verifiable
Dataset Summary
NuminaMath-1.5-RL-Verifiable is a curated subset of the NuminaMath-1.5 dataset, specifically filtered to support reinforcement learning applications requiring verifiable outcomes. This collection consists of 131,063 math word problems from the original dataset that meet strict filtering criteria: all problems have definitive numerical answers, validated problem statements and solutions, and come from… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-RL-Verifiable.verifiable-code-reasoning
Verifiable Code Reasoning
Execution-verified Python problems with chain-of-thought
Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text
Overview
Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests.
Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if:
a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.VerifiableQFT
Verifiable Synethetic QFT Problems
This dataset provides the synthetic QFT problems and rejection sampled CoT samples used in Fine-Tuning Small Reasoning Models for Quantum Field Theory by N. Woodward et al..
The dataset consists of 2,588 synthetic Quantum Field Theory problems with auto-verifiable code solutions and 24,918 rejection-sampled chain-of-thought (CoT) solutions for supervised fine-tuning.
Dataset Summary
This dataset provides two complementary… See the full description on the dataset page: https://huggingface.co/datasets/nswoodward/VerifiableQFT.verifiable-corpus
verifiable-corpus
This is the corpus from "Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning".
Code: https://github.com/jonhue/ttc
Introduction
We study how large language models (LLMs) can continually improve at reasoning on their target tasks at test-time. We propose an agent that assembles a task-specific curriculum, called test-time curriculum (TTC-RL), and applies reinforcement learning to continue training the model for its target task.… See the full description on the dataset page: https://huggingface.co/datasets/lasgroup/verifiable-corpus.polaris_filtered_nemotron_medium_math_verifiable
Polaris Filtered Nemotron Medium Sympy Verifiable (v2)
This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using math verify-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset.
Dataset Summary
Total Original Samples: 2,424,392
Final Kept Samples: 357,790 (14.8%)
Target Reasoning Length: 4k-8k… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_math_verifiable.polaris_filtered_nemotron_medium_sympy_verifiable
Polaris Filtered Nemotron Medium Sympy Verifiable
This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using sympy-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset.
Dataset Summary
Total Original Samples: 2,500,820
Final Kept Samples: 263,123 (10.5%)
Target Reasoning Length: Optimized for… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_sympy_verifiable.Medical-o1-verifiable-problem-Thai
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-verifiable-problem-Thai.verifiable-pythonic-function-calling-lite-thai
wannaphong/verifiable-pythonic-function-calling-lite-thai
make dataset from https://huggingface.co/datasets/driaforall/verifiable-pythonic-function-calling-lite
polaris_filtered_nemotron_easy_math_verifiable
Polaris-Filtered Nemotron Easy Math (Verifiable)
A filtered subset of nvidia/Nemotron-Math-v2 (low / easy split), retaining only non-TIR samples with verifiable boxed answers.
Filtering Pipeline
Remove TIR / tool-use samples — drop any sample that contains Python code blocks (\``python, <|python_start|>, ) or has a non-empty tools/tool` field.
Polaris decontamination — remove samples whose user prompt shares any 15-gram overlap with problems in… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_easy_math_verifiable.medical-o1-verifiable-problem-mk
Dataset Card for Dataset Name
This is a preview of a Macedonian translation of the medical-o1-verifiable-problem dataset by Freedom Intelligence.
Note that this preview currently contains 1068 rows.
Dataset Details
Dataset Structure
Each example consists of a question and a verifiable answer.
Dataset Creation
For methodological details regarding the creation of the original dataset, please refer to the original paper.
Machine translation was… See the full description on the dataset page: https://huggingface.co/datasets/ilijalichkovski/medical-o1-verifiable-problem-mk.NuminaMath-1.5-RL-Verifiable
Dataset Card for NuminaMath-1.5-RL-Verifiable
Dataset Summary
NuminaMath-1.5-RL-Verifiable is a curated subset of the NuminaMath-1.5 dataset, specifically filtered to support reinforcement learning applications requiring verifiable outcomes. This collection consists of 131,063 math word problems from the original dataset that meet strict filtering criteria: all problems have definitive numerical answers, validated problem statements and solutions, and come from… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/NuminaMath-1.5-RL-Verifiable.spai-ss6-corpus-medical-o1-verifiable
SPAI SS6 Medical O1 Verifiable Thai Index
Index repo for the imported Thai medical verifiable-problem dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: medical_o1_verifiable_problem_thai
Rows in canonical config: 40,906
Parquet size in canonical config: 0.00 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-o1-verifiable.Skywork-OR1-Math-Verifiable-Dedup
Skywork OR1 Math — Verifiable, Deduplicated
97,809 questions. Only parser-compatible references are included. Here, “verifiable” means every reference component parses with Math-Verify 0.8.0, with string fallback disabled. It does not mean that the answer has been independently proved correct or that grading model outputs is error-free.
Skywork was processed independently: retain its math rows, clean them, deduplicate within Skywork, screen benchmark overlap and prompt leakage… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/Skywork-OR1-Math-Verifiable-Dedup.DeepScaleR-Verifiable-Dedup
DeepScaleR — Verifiable, Deduplicated
37,713 questions. Only parser-compatible references are included. Here, “verifiable” means every reference component parses with Math-Verify 0.8.0, with string fallback disabled. It does not mean that the answer has been independently proved correct or that grading model outputs is error-free.
DeepScaleR was processed independently: clean its questions, deduplicate within DeepScaleR, screen benchmark overlap and prompt leakage, then retain… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/DeepScaleR-Verifiable-Dedup.Skywork-DeepScaleR-Merged-Verifiable-Dedup
Skywork + DeepScaleR — Verifiable, Cross-Deduplicated
98,941 questions. Only parser-compatible references are included. Here, “verifiable” means every reference component parses with Math-Verify 0.8.0, with string fallback disabled. It does not mean that the answer has been independently proved correct or that grading model outputs is error-free.
The two independently cleaned pools were merged, cross-source duplicate questions were collapsed, unresolved cross-source answer… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/Skywork-DeepScaleR-Merged-Verifiable-Dedup.
