datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Bench-Verified-O1-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.Lego-RL-SWE-Bench-Verified
Lego-RL-SWE-Bench-Verified
The 500 SWE-bench Verified instances as ready-to-run harbor RL environments —
the exact evaluation set behind every SWE-bench Verified number in
LEGO-RL, packaged the same way as the training
set Lego-X/Lego-RL-2699 so
one trainer reads both.
Two parallel views of the same 500 instances:
View
Path
What it is
Official SWE-bench records
swebench_verified_official_500/
The upstream princeton-nlp/SWE-bench_Verified rows, verbatim
Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.agent-trajectories-swe-bench-test-minus-verified
Agent Trajectories: SWE-bench Test \ Verified — Mixed Teachers (gpt-5.2 / gpt-5-mini)
Summary
Full multi-turn agent trajectories collected from the SWE-bench Test minus Verified split
(i.e., SWE-bench Test instances that are not part of SWE-bench Verified).
Intended for SFT of agent models on coding tasks.
Data Collection
Each trajectory was produced by a GT-aware lookahead agent that, at every turn:
Sampled a candidate response from both gpt-5.2 and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/agent-trajectories-swe-bench-test-minus-verified.SWE-Bench_Verified_ABS
SWE-Bench_Verified_ABS
A dataset of 500 software engineering instances derived from SWE-bench,
extended with model-generated test patches for evaluation.
GitHub: OpenAgentEval/SWE-ABS
Dataset Description
Each instance corresponds to a real GitHub issue and pull request.
The dataset includes the original SWE-bench fields. test_patch is replaced
with a model-generated test patch, and the original is preserved as original_test_patch.
Fields
Fields inherited from… See the full description on the dataset page: https://huggingface.co/datasets/OpenAgentLab/SWE-Bench_Verified_ABS.SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920
M2.7 Solo and self-orchestration: three independent eval150 runs each
All six fresh runs completed the same150 tasks and passed original result/trajectory/task/attempt/fingerprint audits. No previous scores were pooled. Each repeat starts new model processes and cold KV caches after a real telemetry smoke. True failed tasks are retained; infrastructure retries are preserved separately.
Mode
Repeat1
Repeat2
Repeat3
Mean /150
Sample SD
m27-solo
94
98
90
94.00
4.00… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920.SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.SWEbench-Verified-eval150-M2.7-OPD143-appendonly-ablation-2repeats-w32-20260920
M2.7 + mixed-OPD143: coordinator append-only ablation
Two new append-only eval150 repeats versus the two existing original-harness repeats (89 and85/150). Same held-out150, models and32-way concurrency; controls were run earlier, not simultaneously. Original controls are reused without rerunning or pooling.
Coordinator history
Repeat1
Repeat2
Mean accuracy
Mean full150 min
original
89
85
58.00%
50.77
append_only
80
90
56.67%
110.30
Worker: mixed-OPD… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD143-appendonly-ablation-2repeats-w32-20260920.SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921
OPD149 bounded evidence7168→3072, latest state at tail: ONE eval150, three 50-task shards
82/150 (54.67%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 69.28 minutes, including queue and any recovery gaps, excluding prior service startup/smoke.
Original two-message coordinator/system prompt; only originally visible complete evidence… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921.devstral-swe-bench-verified
Devstral SWE-bench Verified Trajectories
This dataset contains 4,000 mini-SWE-agent trajectories generated by
mistralai/Devstral-Small-2-24B-Instruct-2512 on a fixed 250-problem training split of
SWE-bench Verified. There are 16 independently seeded rollouts per problem.
Every row preserves the exact request prefix mini-SWE-agent supplied before Devstral's
first generation, the complete ordered rollout, a sanitized final patch, and termination
metadata. Official SWE-bench… See the full description on the dataset page: https://huggingface.co/datasets/suryadv/devstral-swe-bench-verified.SWEbench-Verified-eval150-M2.7-new-checkpoints-appendonly-w64-20260920
New coding checkpoints: append-only M2.7 + 9B, 64 episodes
Four independent full150 screenings, one repetition each. Historical original-harness and 32-concurrency results are not matched controls; no matched baseline delta or stable gain is claimed.
Training arm
Checkpoint
Resolved /150
Accuracy
Full150 min
Exact model
opd-mix-raw
149 final
93
62.00%
49.92
HF
graded350-opd
149 final
80
53.33%
83.08
HF
graded124-opd
149 final
86
57.33%
54.82
HF
solo350-local63… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-new-checkpoints-appendonly-w64-20260920.SWEbench-Verified-eval150-M2.7-OPD149-visible-history-v2-2repeats-w64-20260921
OPD149 visible-history v2: two repeats,64 episodes; historical v1 reference
Two independent v2 full150 evaluations with OPD149 weights,64 episodes and CPU-quota-aware sandboxes. User requested reusing historical v1 score93/150 rather than a new control run. Historical v1 predates the CPU environment repair and is noncontemporaneous, so this is screening evidence, not a matched causal comparison or proof of non-inferiority.
Training arm
Checkpoint
Resolved /150
Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-visible-history-v2-2repeats-w64-20260921.SWE-bench_Verified
SWE-bench Verified - Random Subset (100 instances)
This is a randomly selected subset of 100 instances from princeton-nlp/SWE-bench_Verified.
Dataset Details
Source: princeton-nlp/SWE-bench_Verified (split: test)
Subset Size: 100 instances
Selection Method: Random sampling
Random Seed: 42
Created: Automatically generated
Instance IDs
The following instances are included in this subset:
astropy__astropy-13398
astropy__astropy-14508
astropy__astropy-14539… See the full description on the dataset page: https://huggingface.co/datasets/jerry128/SWE-bench_Verified.SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921
OPD149 evidence retained, latest state at tail: ONE eval150, three 50-task shards
86/150 (57.33%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 71.76 minutes, including queue and any recovery gaps, excluding prior service startup/smoke.
Original two-message coordinator/system prompt; newly exposed evidence events retained, current… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921.swebench-verified-kimi-k2p6-traces
SWE-bench Verified Kimi K2.6 Reasoning Traces
This dataset contains reasoning traces generated on princeton-nlp/SWE-bench_Verified using fireworks_ai/kimi-k2p6-high with a mini-swe-agent based harness. It is intended for research and distillation of software-engineering agents.
The repository is published with three configs because each table has a different schema:
raw_trajectories: one row per SWE-bench instance with the patch, sanitized result JSON, full trajectory JSON, message… See the full description on the dataset page: https://huggingface.co/datasets/MemoryAsModality/swebench-verified-kimi-k2p6-traces.SWE-Bench-Verified-Quick
SWE-Bench-Verified-Quick
Quick-eval subset of
princeton-nlp/SWE-bench_Verified
(SWE-bench paper): 468 / 500 Verified instances.
Default dataset of the swebench_v1 taskset; also supported by the v0 mini_swe_agent_plus
environment.
Changes vs upstream
Latency subset only — drops the slowest-running instances so a full benchmark pass
finishes in under ~30 minutes at reasonable concurrency. No validation semantics; "Verified"
in the name is OpenAI's human… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Bench-Verified-Quick.
