PrimeIntellect/SWE-rebench-V2-Filtered-Easy-Verified
SWE-rebench-V2-Filtered-Easy-Verified Easy slice of PrimeIntellect/SWE-rebench-V2-Filtered-Verified: rows whose upstream LLM-judge difficulty is easy (implementation-time estimate < 15 min). Useful as a lower-variance starting pool for RL curricula. Changes vs upstream Pure slice of the Filtered-Verified set — it inherits every filter and verification pass from the parent (see its card), including the pass-2 flaky removal, no-edit pass, and repo/image blocklists… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-rebench-V2-Filtered-Easy-Verified.
SWE-rebench-V2-Filtered-Easy-Verified

Easy slice of `PrimeIntellect/SWE-rebench-V2-Filtered-Verified`: rows whose upstream LLM-judge difficulty is easy (implementation-time estimate < 15 min). Useful as a lower-variance starting pool for RL curricula.
Changes vs upstream
- Pure slice of the Filtered-Verified set — it inherits every filter and verification pass from the parent (see its card), including the pass-2 flaky removal, no-edit pass, and repo/image blocklists that the earlier standalone
-Easy-Cleanderivation lacked. image_namecarries the raw upstream Docker Hub source ref (docker.io/swerebenchv2/<name>:<tag>): since the platform's 2026-07-15 org-less image migration (ENG-4518), source refs resolve natively on Prime, so no registry rewrite is needed. (The parent still temporarily carriesprime/primeintellect/...refs for prod-training compatibility; any such ref is mapped back to its source form here.)
License mirrors upstream: CC-BY-4.0.
Splits
How to use
Install the `swerebench_v2_v1` taskset from research-environments, then run it end-to-end with verifiers:
uv pip install --prerelease=allow "git+https://github.com/PrimeIntellect-ai/research-environments.git#subdirectory=environments/swe/swerebench_v2_v1"
uv run eval --taskset.id swerebench_v2_v1 -m <your-model> -n 100 -r 4Generation
<details> <summary>Reproduction script — <code>swe-rebench-v2-filtered-easy-verified.py</code></summary>
This dataset was created by running:
uv run datasets/swe-rebench-v2-filtered-easy-verified.py -H# swe-rebench-v2-filtered-easy-verified.py
"""Derive the easy slice of `PrimeIntellect/SWE-rebench-V2-Filtered-Verified`.
Rows whose upstream LLM-judge ``meta.llm_metadata.difficulty`` is ``"easy"``
(implementation-time estimate < 15 min). A pure slice of the already
filtered-and-verified parent, so it inherits every gate from
``swe-rebench-v2-filtered-verified.py`` — including the pass-2 flaky
removal, the no-edit pass, and the repo/image blocklists that the earlier
standalone ``SWE-rebench-V2-Easy-Clean`` derivation (prime-data PR #21,
closed unmerged) lacked.
One transform on top of the slice: ``image_name`` is mapped back to the raw
upstream Docker Hub source ref (``prime/primeintellect/<name>:<tag>`` →
``docker.io/swerebenchv2/<name>:<tag>``). Since the platform's 2026-07-15
org-less image migration (ENG-4518), source refs resolve natively on Prime,
so no registry rewrite is needed. The parent dataset still temporarily
carries the ``prime/primeintellect/`` rewrite for prod-training
compatibility (until its 6,275 images finish transferring to the platform
registry); once the parent is reverted to source refs too, this mapping
becomes a no-op.
"""
# /// script
# requires-python = ">=3.12"
# dependencies = ["datasets>=4.0.0", "jinja2"]
# ///
import argparse
import sys
from pathlib import Path
from typing import cast
from huggingface_hub import create_repo, whoami
from datasets import Dataset, load_dataset
SOURCE_REPO = "PrimeIntellect/SWE-rebench-V2-Filtered-Verified"
_PRIME_PREFIX = "prime/primeintellect/"
_SOURCE_PREFIX = "docker.io/swerebenchv2/"
def _restore_source_image_ref(image_name: str) -> str:
# Undo the parent's prime-registry rewrite; no-op for refs already in
# source form (i.e. once the parent itself is reverted to source refs).
if image_name.startswith(_PRIME_PREFIX):
return _SOURCE_PREFIX + image_name[len(_PRIME_PREFIX) :]
return image_name
def _swe_card(key: str):
"""Build this dataset's card from the shared SWE card registry (swe_cards.py)."""
sys.path.insert(0, str(Path(__file__).resolve().parent))
from swe_cards import build_card
return build_card(key)
def _is_easy(example: dict) -> bool:
llm = (example.get("meta") or {}).get("llm_metadata") or {}
return llm.get("difficulty") == "easy"
def prepare_data(source_repo: str) -> Dataset:
ds = cast(Dataset, load_dataset(source_repo, split="train"))
n_before = len(ds)
easy = ds.filter(_is_easy, num_proc=8)
print(f"Kept {len(easy):,} / {n_before:,} rows with difficulty == easy")
return easy.map(
lambda ex: {"image_name": _restore_source_image_ref(ex["image_name"])},
num_proc=8,
load_from_cache_file=False,
)
def main(repo_name: str, push_to_hub: bool, private: bool, source_repo: str) -> None:
print(f"⚙️ Slicing {source_repo} to difficulty == easy")
dataset = prepare_data(source_repo)
card = _swe_card("swe-rebench-v2-filtered-easy-verified")
if push_to_hub:
create_repo(repo_name, private=private, repo_type="dataset", exist_ok=True)
card.push_to_hub(repo_name, repo_type="dataset")
dataset.push_to_hub(repo_name, private=private)
print(f"✅ Pushed dataset to https://huggingface.co/datasets/{repo_name}")
else:
print("ℹ️ Skipped pushing to HF Hub. To push, use the `--push-to-hub` or `-H` flag.")
def check_write_access(org: str):
is_authed = False
try:
info = whoami()
token = info["auth"]["accessToken"]["displayName"]
for entity in info["auth"]["accessToken"]["fineGrained"]["scoped"]:
if entity["entity"]["name"] == org and "repo.write" in entity["permissions"]:
is_authed = True
except Exception:
raise ValueError("❌ You are not logged in. Please run `hf auth login` or `export HF_TOKEN=...`")
if not is_authed:
raise ValueError(f"❌ Your current token `{token}` does not have write access to `{org}`")
print(f"✅ Confirmed write access with token `{token}` to `{org}`")
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument(
"--username", "-U", default="PrimeIntellect", type=str, help="The username to push the dataset to."
)
parser.add_argument(
"--dataset-name",
"-D",
default="SWE-rebench-V2-Filtered-Easy-Verified",
type=str,
help="The dataset name.",
)
parser.add_argument("--source-repo", "-S", default=SOURCE_REPO, type=str, help="The parent dataset to slice.")
parser.add_argument("--push-to-hub", "-H", action="store_true", help="Whether to push the dataset to the hub.")
parser.add_argument("--dataset-private", "-p", action="store_true", help="Whether to make the dataset private.")
args = parser.parse_args()
assert len(args.dataset_name.split("/")) == 1, "Dataset name must not include the username"
if args.push_to_hub:
check_write_access(args.username)
main(
repo_name=f"{args.username}/{args.dataset_name}",
push_to_hub=args.push_to_hub,
private=args.dataset_private,
source_repo=args.source_repo,
)
</details>
Original Dataset Card
Snapshot of the `nebius/SWE-rebench-V2` card at card-build time — see the live card for updates.
<details> <summary>Original <code>nebius/SWE-rebench-V2</code> dataset card</summary>
SWE-rebench-V2
Dataset Summary
SWE-rebench-V2 is a curated dataset of software-engineering tasks derived from real GitHub issues and pull requests. The dataset contains 32,079 samples covering Python, Go, TypeScript, JavaScript, Rust, Java, PHP, Kotlin, Julia, Elixir, Scala, Swift, Dart, C, C++, C#, R, Clojure, OCaml, and Lua. For log parser functions, base Dockerfiles, and the prompts used, please see https://github.com/SWE-rebench/SWE-rebench-V2 The detailed technical report is available at “SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale”.
Quick Start
from datasets import load_dataset
ds = load_dataset("nebius/SWE-rebench-V2", split="train")
print(len(ds)) # 32079Dataset Structure
License
The dataset is licensed under the Creative Commons Attribution 4.0 license. However, please respect the license of each specific repository on which a particular instance is based. To facilitate this, the license of each repository at the time of the commit is provided for every instance.
Citation
@misc{badertdinov2026swerebenchv2languageagnosticswe,
title={SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale},
author={Ibragim Badertdinov and Maksim Nekrashevich and Anton Shevtsov and Alexander Golubev},
year={2026},
eprint={2602.23866},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2602.23866},
}
</details>
