datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GRPO_Responsechartqa-grpodoclaynet-grposciegqa-grpo
SciEGQA-GRPO: Single-Page Evidence Grounded QA Dataset
Converted from SciEGQA-Train for GRPO reinforcement learning training on the SLIME framework.
Filter
Only kept samples where evidence is on a single page (multi-page evidence samples removed)
Original: 30,780 samples → Filtered: 19,180 samples
Format
File: train.jsonl (22 GB)
Each line is a JSON object with the following fields:
Field
Type
Description
problem
str
Prompt for QA task:… See the full description on the dataset page: https://huggingface.co/datasets/PassionPrc/sciegqa-grpo.grpo-qwen1.5b-textworld-policy-logitsdocvqa-grpoStaleness-GRPO-DAPO-Math-17k
Staleness GRPO DAPO Math 17k
The exact 17,005-row training dataset shared by the staleness-cap-2 Qwen2.5-Math-1.5B, Qwen2.5-3B, and Qwen2.5-Math-7B checkpoints, and the staleness-cap-4 Qwen2.5-Math-1.5B checkpoint. All four training manifests record the same SHA-256 for the training file.
Source and processing
Derived from the all configuration of open-r1/DAPO-Math-17k-Processed, itself processed from BytedTsinghua-SIA/DAPO-Math-17k. Source revision:… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/Staleness-GRPO-DAPO-Math-17k.funsd-grpoManim-grpo-dataset-200
Manim GRPO Dataset 200
200+ cleaned ManimGL scene excerpts and populated metadata bundles for GRPO / reward-model training on mathematical animation code. Each problem is a directory data/problems/MB-XXX/ containing reference.py extracted from 3b1b/videos (years 2022–2026), complete with problem.json, visual_events.json, coverage.json, version_notes.json, and ref_embeddings.npy.
Dataset structure
data/
problems/
MB-001/ … MB-200/
reference.py… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/Manim-grpo-dataset-200.GRPO-Lambda-ParsedForUnsloth
This is a dataset I've generated containing about 50,000 JSON datapoints of the lambda calculus, including prompts for the problem, and extended explanation of the solution process. It was originally generated in pure JSON with extra metadata for myself. There are 50 separate files ranging on average 2.7 MB, each containing 1000 datapoints, this is split for the at home trainers with low GPU. This dataset is converted to a Question-Reasoning-Answer style ideally (hopefully) for use with… See the full description on the dataset page: https://huggingface.co/datasets/Creekside/GRPO-Lambda-ParsedForUnsloth.Go-GRPO-1K
Go-GRPO-1K
Paper | Code
Project Context
The LoGos model uses this dataset to transfer reasoning capabilities acquired from long CoT data to Go tasks. Through mixed fine-tuning and reinforcement learning, the model learns to perform analysis, reasoning, and summarization to select optimal moves on the Go board.
Citation
If you find this dataset useful for your research, please cite:
@misc{ma2026mixingexpertknowledgebring,
title={Mixing Expert Knowledge:… See the full description on the dataset page: https://huggingface.co/datasets/YichuanMa/Go-GRPO-1K.FinRAG-GRPO
FinRAG-GRPO Preference Dataset
A Chinese-language preference dataset for training Reasoning Reward Models (ReasRM) via GRPO-based reinforcement learning.
🚧 This dataset is actively maintained and will be expanded with additional domains and languages over time.
Dataset Summary
This dataset contains pairwise preference samples designed to train a reward model that reasons before judging — the model generates an evaluation rationale before outputting a preference label… See the full description on the dataset page: https://huggingface.co/datasets/SamWang0405/FinRAG-GRPO.webgen-agent_train_step-grpo
WebGen-Agent
WebGen-Agent is an advanced website generation agent designed to autonomously create websites from natural language instructions. It was introduced in the paper WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning.
Code: https://github.com/mnluzimu/WebGen-Agent
Project Overview
WebGen-Agent combines state-of-the-art language models with specialized training techniques to create a powerful… See the full description on the dataset page: https://huggingface.co/datasets/luzimu/webgen-agent_train_step-grpo.funding-extraction-artifact-data-mix-grpo-mixed-reward
Funding Extraction Training Data
Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements.
Dataset Structure
data/
├── full/ # Complete unsplit dataset
│ ├── train.jsonl # 5,264 real Crossref funding statements
│ └── synthetic.jsonl # 10,124 LLM-generated funding statements
├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.BIRD-GRPO-5Kbash-agent-grpo-pairs
Bash Agent GRPO Pairs
Single-turn (intent → shell command) pairs for training a small, local, Claude-Code-style
bash agent with SFT or GRPO. Each record pairs a natural-language objective with exactly
one verifiable bash command, framed as a single-tool bash(command, description) call.
The dataset is designed to be rewardable: the ground-truth command is a deterministic target,
so a shell-equivalence reward (canonical program + flag set + argument comparison) can score… See the full description on the dataset page: https://huggingface.co/datasets/adeelahmad/bash-agent-grpo-pairs.JOSIE-Zero-4-GRPO-Outputs-105
JOSIE-Zero-4: Out-of-Training GRPO Evaluation Generations
This dataset contains 105 model-generated reasoning traces from JOSIE-Zero-4. The samples were selected from a 500-example evaluation on the Math branch of openbmb/UltraData-RL-2609.
These problems were not used to train JOSIE-Zero-4 with GRPO. The evaluation was designed to examine whether reasoning behavior learned through GRPO transfers to data outside the model's training distribution.
Of the 500 evaluated generations… See the full description on the dataset page: https://huggingface.co/datasets/Goekdeniz-Guelmez/JOSIE-Zero-4-GRPO-Outputs-105.minimind-stage2-grpoamazon-grpo-ssmrm8488__phi-4-14B-grpo-gsm8k-3e-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-gsm8k-3e
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-gsm8k-3e
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-gsm8k-3e-details.Codeforces-Cleaned-GRPO
Codeforces-Cleaned-GRPO
Deep-cleaned for GRPO/RL training | 500 examples | 30 bugs fixed
📋 Dataset Description
Competitive programming problems from Codeforces via Open-R1. This cleaned version removes HTML math notation tags and whitespace issues that would break problem descriptions during training.
Original source: open-r1/codeforces by Open-R1
📊 Cleaning Statistics
Metric
Value
Original examples
500
Clean examples
500… See the full description on the dataset page: https://huggingface.co/datasets/Eyght/Codeforces-Cleaned-GRPO.Tiny-Maze-Mock-GRPOshotpath-grpo-trajectory-audit-20260729
ShotPath GRPO trajectory audit
This audit reconstructs the committed 260-step trajectory by keeping the last logged occurrence of each pair after time-limit rollbacks. It includes compact statistics for every committed group and detailed candidate/judge records plus pre/post images for stratified and contrast samples.
GRPO-Reasoning-Tools-Cleaned
GRPO-Reasoning-Tools-Cleaned
Deep-cleaned for GRPO/RL training | 1,998 examples | 2,608 bugs fixed
📋 Dataset Description
Structured reasoning and tool-use prompts designed for GRPO training. This cleaned version normalizes whitespace, removes XML reasoning tags, and filters non-English content.
Original source: nphearum/grpo-4k-reasoning-tools by Independent
📊 Cleaning Statistics
Metric
Value
Original examples
2,000
Clean… See the full description on the dataset page: https://huggingface.co/datasets/Eyght/GRPO-Reasoning-Tools-Cleaned.xfund-grpomanibench-grpo
ManiBench GRPO Reference Scenes
200 cleaned ManimGL scene excerpts for GRPO / reward-model work on math animation code. Each problem is a folder data/problems/MB-XXX/ with a reference.py extracted from 3b1b/videos (years 2022–2026).
This release is reference code only. Prompt, visual-event, coverage, and version-note JSON files are empty placeholders to fill later. CLIP embeddings and raw video are not included.
Not in this set: the 12 ManiBench pilot / benchmark videos… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/manibench-grpo.DeepSeek-Reasoning-Cleaned-GRPO
DeepSeek-Reasoning-Cleaned-GRPO
Deep-cleaned for GRPO/RL training | 500 examples | 41 bugs fixed
📋 Dataset Description
Pure reasoning traces from DeepSeek R1 distilled mathematical reasoning data. This cleaned version removes markdown headers, control characters, and whitespace that would confuse model training.
Original source: open-r1/OpenR1-Math-220k by DeepSeek/Open-R1
📊 Cleaning Statistics
Metric
Value
Original examples
500… See the full description on the dataset page: https://huggingface.co/datasets/Eyght/DeepSeek-Reasoning-Cleaned-GRPO.scugnizz-grpo-chain-v1Dongwei__DeepSeek-R1-Distill-Qwen-7B-GRPO-details
Dataset Card for Evaluation run of Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
Dataset automatically created during the evaluation run of model Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dongwei__DeepSeek-R1-Distill-Qwen-7B-GRPO-details.mrm8488__phi-4-14B-grpo-limo-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-limo
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-limo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-limo-details.
