datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Bench-Pro-interactive-issue-qa
SWE-Bench Pro / Interactive / Issue+QA
Companion data release for the anonymous paper "Opinion: Coding-Agent Benchmarks Should Match Their Users' Task Flows" (SWE-TaskFlow). This dataset contains the QA-augmented trajectories: verifiable questions about repository behavior inserted before, between, or after the split issue turns. Every question ships with its hidden reference answer, an executable golden proof script, and the creation-time proof-execution report.
The plain… See the full description on the dataset page: https://huggingface.co/datasets/Anonym01048/SWE-Bench-Pro-interactive-issue-qa.SWE-Bench-Pro-interactive-issue
SWE-Bench Pro / Interactive / Issue
Companion data release for the anonymous paper "Opinion: Coding-Agent Benchmarks Should Match Their Users' Task Flows" (SWE-TaskFlow). SWE-TaskFlow transforms an issue-derived benchmark into replayable multi-turn trajectories while preserving the original tasks and tests. This dataset contains the issue-solving prompt sequences (no QA turns) over the 701 SWE-Bench Pro tasks that admit a three-part decomposition (out of the 731 public tasks).… See the full description on the dataset page: https://huggingface.co/datasets/Anonym01048/SWE-Bench-Pro-interactive-issue.
