datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
R2E-Gym-LiteR2E-Gym-V1R2E-Gym-SubsetR2E-Gym-Subset-Verified
R2E-Gym-Subset-Verified
Gold-patch-validated subset of
R2E-Gym/R2E-Gym-Subset
(paper). The train split contains
4,522 / 4,578 rows (98.78%) verified scoreable end-to-end: apply the gold patch, run the
upstream /testbed/run_tests.sh baked into the row's image, check the parsed outcomes against
expected_output_json.
Changes vs upstream
Validation-only subset — our passes, run in fresh sandboxes per row: one full pass at
concurrency 200, then a 10× retry pass over… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/R2E-Gym-Subset-Verified.R2E-Gym-Full
R2E-Gym Subset Filtered for MAGRPO
Filtered subset of R2E-Gym optimized for 2-agent MAGRPO training with 7B models.
Dataset Statistics
Total instances: 167
Format: Issue description + Oracle files in prompt
Optimized for: 2-agent collaboration, 7B models
Filtering Criteria (SWE-bench Lite Style)
Problem statement: >40 words (up to 500 for context window)
Must have non-empty oracle patch (non-test file changes)
File count: Exactly 1 oracle file (single-file… See the full description on the dataset page: https://huggingface.co/datasets/ryankamiri/R2E-Gym-Full.R2E-Gym-CollabR2E-Gym-Subset-OracleR2E-Gym-SubsetR2E-Gym-Lite-Truncate-7B-FixedR2E-Gym-Subset-contextR2E-Gym
R2E-Gym: OpenCode, Codex, and Claude Code environments
This repository contains three environment variants of
R2E-Gym/R2E-Gym-Subset.
Each variant contains the same 4,578 tasks in five Parquet shards, with the
original 14-column schema and task order. Only docker_image is replaced with
the corresponding public image containing the selected agent.
Configuration
Files
Docker image prefix
opencode
opencode/data/*.parquet
docker.io/loongsage/r2e-gym:opencode_
codex… See the full description on the dataset page: https://huggingface.co/datasets/loongsage/R2E-Gym.R2E-Gym-Lite-Truncate-7BR2E-Gym-Yiming-Combined-1p6pctR2E-Gym-Yiming-Combined-0p8pctR2E-Gym-Subset
R2E-Gym Subset Filtered for MAGRPO
Filtered subset of R2E-Gym optimized for 2-agent MAGRPO training with 7B models.
Dataset Statistics
Total instances: 54
Format: Issue description + Oracle files in prompt
Optimized for: 2-agent collaboration, 7B models
Filtering Criteria (SWE-bench Lite Style)
Problem statement: >40 words (up to 500 for context window)
Must have non-empty oracle patch (non-test file changes)
File count: Exactly 1 oracle file (single-file… See the full description on the dataset page: https://huggingface.co/datasets/ryankamiri/R2E-Gym-Subset.R2E-Gym-Lite-Truncate-HeuristicR2E-Gym-Subset-rllmR2E-Gym-TrajsR2E-Gym-Lite-with-DifficultyR2E-Gym-Yiming-Combined-20pctR2E-Gym-Yiming-From-Clean-1p6pct-4274R2E-Gym-Javi-40pct-4274-NoTestR2E-Gym-Poisoned-Yiming-Distribution-Matched-20-3kR2E-Gym-Yiming-Combined-30pctR2E-Gym-Yiming-From-Clean-0p8pct-4274R2E-Gym-Yiming-From-Clean-10pct-4274R2E-Gym-Javi-Hard-40pct-4274-InstallTestR2E-Gym-Yiming-From-Clean-50pct-4274R2E-Gym-Lite-Difficulty-Clean-OnlyR2E-Gym-Javi-From-Clean-60pct-4274
