datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swepro-luna-matched-pair
SWE-bench Pro, matched pair: Ouroboros vs Codex CLI on one model
Status: Self-reported matched-pair study. Both harnesses used the same
model, task set and evaluator. The strict result is a statistical tie.
Start here
Strict result
Ouroboros 58.2%, Codex CLI 59.4%, McNemar p = 0.40
Model
openai/gpt-5.6-luna for both arms
Filter
655 paired tasks after the same reference-leak filter was applied to both arms
Exact evidence
6228037, manifest.csv… See the full description on the dataset page: https://huggingface.co/datasets/razzant/swepro-luna-matched-pair.Monthly-SWEBench-2026-05
Monthly-SWEBench 2026-05
This package contains the 2026-05 Monthly-SWEBench final release set. It includes 100 Harbor-format software engineering tasks selected from closed GitHub PRs and validated with oracle=1 / nop=0.
Files
bugfix.tar.zst: 50 bug-oriented repair or maintenance tasks.
non_bugfix.tar.zst: 50 feature, API evolution, or engineering-improvement tasks.
preview.csv: task ids, split labels, source change buckets, and archive paths.
tasks.conf: one… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-05.SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.Monthly-SWEBench-2026-03
Monthly-SWEBench-2026-03
Monthly-SWEBench-2026-03 is a curated benchmark of 112 real-world software engineering tasks, sourced from GitHub pull requests merged in March 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent.
View leaderboard and results →
112 tasks — 68 bugfix + 44 non-bugfix
Tasks span diverse open-source repositories
Each task includes a runnable environment, test suite, and reference solution
Task Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-03.Monthly-SWEBench-2026-04
Monthly-SWEBench-2026-04
Monthly-SWEBench-2026-04 is a curated benchmark of 90 real-world software engineering tasks, sourced from GitHub pull requests merged in April 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent.
View leaderboard and results →
90 tasks — 43 bugfix + 47 non-bugfix
Tasks span diverse open-source repositories
Each task includes a runnable environment, test suite, and reference solution
Task Structure
Each task is a… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-04.qwen9b-coop-mini-swe-agent
qwen9b-coop-mini-swe-agent
Two-agent cooperative coding trajectories generated by running
CooperBench in coop mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote.
The matched solo version is at
CooperBench/qwen9b-solo-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-mini-swe-agent.qwen9b-solo-mini-swe-agent
qwen9b-solo-mini-swe-agent
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
One agent implements both features in each task.
The matched coop version is at
CooperBench/qwen9b-coop-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.sarcastic-dataset
Sarcastic Dataset
A dataset of 720 sentences with sarcastic and extra-flair versions generated using OpenAI GPT models.
sentence: Original sentence
translation: Sarcastic version
translation_extra: Sarcastic version with extra flair
