datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-SWE-bench
Multi-SWE-bench
Re-upload of ByteDance's Multi-SWE-bench
evaluation benchmark: 2,132 issue-resolving tasks across the seven Multi-SWE languages.
This is the held-out eval benchmark; for RL training data use
PrimeIntellect/Multi-SWE-RL-Verified.
Changes vs upstream
Storage schema only: per-test maps are stored as columnar struct-of-lists so the rows load
cleanly with datasets. Row content is unchanged.
License mirrors upstream: ByteDance licenses the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-bench.fineweb-edu
Pre-shuffled fineweb-edu dataset
Multi-SWE-RL-Verified
Multi-SWE-RL-Verified
Gold-patch-validated subset of
PrimeIntellect/Multi-SWE-RL-Reupload
(ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across
C, Go, Java, JavaScript, Rust, and TypeScript that produce a clean reward signal end-to-end.
Default dataset of the multiswe_v1 taskset.
Changes vs upstream
Starting from the 4,703-row re-upload:
C++ dropped wholesale — 0/449 rows passed gold-patch validation in pass 1; the images are
broken for scoring, not merely… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Verified.Reverse-Text-RL
Reverse-Text-RL
A small, scrappy RL dataset used in prime-rl's CI to debug RL training asking a model to reverse small sentences character-by-character. Follows the general format of PrimeIntellect/Reverse-Text-SFT
The following script was used to generate the dataset.
from datasets import Dataset, load_dataset
dataset = load_dataset("willcb/R1-reverse-wikipedia-paragraphs-v1-1000", split="train")
prompt = "Reverse the text character-by-character. Put your answer in… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Reverse-Text-RL.PrimeVul
PrimeVul
Mirror of the PrimeVul dataset.
StackV1-popular
Stack V1 with popular programming languages
Javascript
Python
C
C++
SQL
Cuda
INTELLECT-3-RLScale-SWE-Verified
Scale-SWE-Verified
Gold-patch-validated fork of
AweAI-Team/Scale-SWE
(paper): 17,202 / 20,181 Python issue-resolving tasks
that produce a clean reward signal end-to-end. Default dataset of the scaleswe_v1 taskset.
Changes vs upstream
Validation (ours) removed 2,979 / 20,181 rows (14.8%):
892 rows whose image_url appears in
scale-swe-exclude-images.json.
2,061 rows categorized gold_patch_failure in
scale-swe-validation.jsonl.
15 rows categorized noop_pass… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Scale-SWE-Verified.SWE-rebench-V2
SWE-rebench-V2
Full re-upload of Nebius's
SWE-rebench-V2
(paper): 32,076 / 32,079 freshly-mined GitHub PR tasks
across 17 languages. Unfiltered mirror for large-scale runs; the curated RL subsets are
SWE-rebench-V2-Filtered-Verified
and
SWE-rebench-V2-Filtered-Easy-Verified.
Changes vs upstream
Dropped exactly 3 rows whose Docker Hub image no longer exists upstream (dhi-mikeio tags
227-50037f2, 442-3f8bb64, 690-7882682 — manifests return denied; list ships as… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-rebench-V2.R2E-Gym-Subset-Verified
R2E-Gym-Subset-Verified
Gold-patch-validated subset of
R2E-Gym/R2E-Gym-Subset
(paper). The train split contains
4,522 / 4,578 rows (98.78%) verified scoreable end-to-end: apply the gold patch, run the
upstream /testbed/run_tests.sh baked into the row's image, check the parsed outcomes against
expected_output_json.
Changes vs upstream
Validation-only subset — our passes, run in fresh sandboxes per row: one full pass at
concurrency 200, then a 10× retry pass over… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/R2E-Gym-Subset-Verified.verifiable-coding-problems
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
PrimeBench
PrimeBench
Practical Real-world Industry and Multi-domain Evaluation benchmark.
PrimeBench is a benchmark for evaluating evaluators. Each of its 400 examples is a pair of
responses to the same prompt, deliberately edited so that one is better than the other along a
named criterion. A reward model or LLM judge passes an example if it scores the chosen response
above the rejected one.
Built and maintained by Composo.
Why it exists
Most preference datasets score… See the full description on the dataset page: https://huggingface.co/datasets/ComposoAI/PrimeBench.zkaedi-prime-constitutions
ZKAEDI PRIME Constitutions
Generated by Qwen2.5-72B-Instruct-AWQ via ZKAEDI PRIME recursive Hamiltonian field dynamics.
Schema
Column
Type
Description
run_id
str
Run folder name e.g. run_001_20260312_184001
timestamp
str
UTC timestamp
tier
int
1=canonical (η≈0.4,β≈0.1,σ≈0.05) · 2=structured
eta/gamma/beta/sigma/seed
float
PRIME field parameters
v1_tokens/v2_tokens/critique_tokens
int
Generation token counts
field_trace
str
Full 100-step μ/σ/max/min… See the full description on the dataset page: https://huggingface.co/datasets/zkaedi/zkaedi-prime-constitutions.SWE-rebench-V2-Filtered-Verified
SWE-rebench-V2-Filtered-Verified
Filtered and gold-patch-verified subset of Nebius's
SWE-rebench-V2
(paper): 6,272 / 32,079 freshly-mined GitHub PR tasks
across 17 languages. Default dataset of the swerebench_v2_v1 taskset.
Changes vs upstream
Filtered (selection — the bulk of the cut):
Upstream's own per-row LLM-judge metadata: difficulty labeled easy/medium/hard, judge grade
code == "A" (clearly solvable), intent_completeness == "complete", no detected_issues… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-rebench-V2-Filtered-Verified.SYNTHETIC-1
SYNTHETIC-1: Two Million Crowdsourced Reasoning Traces from Deepseek-R1
SYNTHETIC-1 is a reasoning dataset obtained from Deepseek-R1, generated with crowdsourced compute and annotated with diverse verifiers such as LLM judges or symbolic mathematics verifiers. This is the raw version of the dataset, without any filtering for correctness - Filtered datasets specifically for fine-tuning as well as our 7B model can be found in our 🤗 SYNTHETIC-1 Collection.
The dataset consists of the… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-1.SWE-Lego-Real-Data-Verified
SWE-Lego-Real-Data-Verified
Gold-patch-validated subset of
PrimeIntellect/SWE-Lego-Real-Data
(itself a fixed fork of SWE-Lego's real-data split). The
resolved split contains 4,323 / 4,432 rows (97.54%) verified scoreable end-to-end: apply
test_patch, apply the gold patch, run the row's test_cmd in its image, require every
F2P/P2P test to report PASSED.
Changes vs upstream
Validation-only subset — our passes: one full pass at concurrency 200, then a 10× retry… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Lego-Real-Data-Verified.Hendrycks-Math
Hendrycks-Math
Generation
This dataset was created by running
uv run hendrycks-math.py -H -p
# hendrycks-math.py
# /// script
# requires-python = ">=3.12"
# dependencies = ["datasets>=4.0.0", "jinja2"]
# ///
import argparse
import json
import sys
import time
from pathlib import Path
from typing import cast
from huggingface_hub import DatasetCard, DatasetCardData, create_repo, whoami
from datasets import Dataset, load_dataset
def prepare_hendrycks_math() ->… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Hendrycks-Math.Reverse-Text-SFT
Reverse-Text-SFT
A small, scrappy SFT dataset used for warming up a small model (e.g. Qwen/Qwen3-0.6B) for RL training. Contains examples in prompt-completion chat format of reversing 5-20 words of text character-by-character. The raw sentences were processed from willcb/R1-reverse-wikipedia-paragraphs-v1-1000.
The following script was used to generate the dataset.
from datasets import Dataset, load_dataset
dataset = load_dataset("willcb/R1-reverse-wikipedia-paragraphs-v1-1000"… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Reverse-Text-SFT.Multi-SWE-RL-Reupload
Multi-SWE-RL-Reupload
Verbatim re-upload of ByteDance's community-sourced
Multi-SWE-RL
(paper): 4,703 containerized issue-resolving tasks across
C, C++, Go, Java, JavaScript, Rust, and TypeScript.
For training, prefer
PrimeIntellect/Multi-SWE-RL-Verified,
the gold-patch-validated subset of this data.
Changes vs upstream
Storage schema only: per-test maps are stored as columnar struct-of-lists so the rows load
cleanly with datasets (the upstream nested structs… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Reupload.NuminaMath-QwQ-CoT-5M
INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning
INTELLECT-MATH is a 7B parameter model optimized for mathematical reasoning. It was trained in two stages, an SFT stage, in which the model was fine-tuned on verified QwQ outputs, and an RL stage, in which the model was trained using the PRIME-RL recipe.
We demonstrate that the quality of our SFT data can impact the performance and training speed of the RL stage: Due to its… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/NuminaMath-QwQ-CoT-5M.SWE-primereal-world-swe-problems
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
INTELLECT-3-SFTsynthetic-code-understanding
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
Eurus-2-RL-Data
Eurus-2-RL-Data
Links
📜 Paper
📜 Blog
🤗 PRIME Collection
Introduction
Eurus-2-RL-Data is a high-quality RL training dataset of mathematics and coding problems with outcome verifiers (LaTeX answers for math and test cases for coding).
For math, we source from NuminaMath-CoT. The problems span from Chinese high school mathematics to International Mathematical Olympiad competition questions.
For coding, we source from APPS, CodeContests, TACO, and Codeforces.… See the full description on the dataset page: https://huggingface.co/datasets/PRIME-RL/Eurus-2-RL-Data.SYNTHETIC-2-SFT-verified
SYNTHETIC-2
SYNTHETIC-2 is an open reasoning dataset spanning a variety of math, coding and general reasoning tasks along with reasoning traces generated in a collaborative manner. The dataset contains both high quality reasoning traces from Deepseek-R1-0528 ideally suited for SFT, as well as multiple reasoning traces from smaller models which can be used for difficulty estimation.
To read more about our data collection approach, check out our blog post.
We release the following… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-2-SFT-verified.stackexchange-question-answering
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
SYNTHETIC-1-SFT-Data
SYNTHETIC-1: Two Million Crowdsourced Reasoning Traces from Deepseek-R1
SYNTHETIC-1 is a reasoning dataset obtained from Deepseek-R1, generated with crowdsourced compute and annotated with diverse verifiers such as LLM judges or symbolic mathematics verifiers. This is the SFT version of the dataset - the raw data and preference dataset can be found in our 🤗 SYNTHETIC-1 Collection.
The dataset consists of the following tasks and verifiers that were implemented in our library… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-1-SFT-Data.INTELLECT-2-RL-Dataset
INTELLECT-2
INTELLECT-2 is a 32 billion parameter language model trained through a reinforcement learning run leveraging globally distributed, permissionless GPU resources contributed by the community.
The model was trained using prime-rl, a framework designed for distributed asynchronous RL, using GRPO over verifiable rewards along with modifications for improved training stability. For detailed information on our infrastructure and training recipe, see our technical report.… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/INTELLECT-2-RL-Dataset.INTELLECT-2-only-math
