datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-rebench-openhands-trajectories
Dataset Summary
SWE-rebench-OpenHands-Trajectories is a dataset of multi-turn agent trajectories for software engineering tasks, collected
using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.54.0) agent scaffolding.
This dataset captures complete agent execution traces as they attempt to resolve real GitHub issues from
nebius/SWE-rebench.
Each trajectory contains the agent's step-by-step reasoning, actions, and environmental observations.
Metric… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-openhands-trajectories.SWE-Zero-openhands-trajectories
SWE-Zero Trajectories: Execution-free Fine-tuning for Software Engineering Agents
Data Overview
SWE-ZERO Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 318k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SWE-Zero-openhands-trajectories.SWE-Hero-openhands-trajectories
SWE-Hero Trajectories: Execution-based Fine-tuning for Software Engineering Agents
Data Overview
SWE-Hero Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 34k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SWE-Hero-openhands-trajectories.OpenHand-Synth
Dataset Card for OpenHand-Synth
📜 Paper: OpenHand-Synth: A Large-Scale Synthetic Handwriting Dataset for Multimodal Language Models
Sample Images
Image
Ground Truth
Source
Language
CER
JW
02-10-1436
faker-date
por
0.10
0.96
Stephan Thomsen-Johansen
faker-name
dan
0.0
1.0
Le chat mange.
tatoeba
fra
0.0
1.0
Classical musicsoothes me.She took the risk, knowing that shemight lose a lot of money.I could not catcha single word of their talk.In the old days… See the full description on the dataset page: https://huggingface.co/datasets/to-be/OpenHand-Synth.OpenHands-Sampled-TrajectoriesSWE-rebench-code-searchOpenHands-SFT-TrajectoriesSWE-bench-devin-fullOpenHands-Verifier-Trajectoriesharbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.SWE-bench_Verified-locagentSWE-smith-py-code-searchopenhands-feedback
OpenHands Feedback Dataset 🙌
Dataset Description
What is OpenHands Feedback?
The OpenHands Feedback Dataset is a collection of user interactions and feedback with the OpenHands AI coding assistant. This dataset contains real-world examples of how users interact with AI coding assistants, including both successful and unsuccessful interactions, along with user feedback on the quality and helpfulness of the responses.
The dataset currently contains 275 examples… See the full description on the dataset page: https://huggingface.co/datasets/OpenHands/openhands-feedback.SWE-bench-devin-passedSWE-bench-devin-full-filteredharbor-devel-sandboxes_glm_4.6_traces_openhandsCodeScout_Eval_RolloutsSWE-Gym-code-searchDevin-SWE-bench-outputopenhands-index
OpenHands Index — Leaderboard Snapshot
Auto-published from https://github.com/OpenHands/openhands-index-results
on every push to main. Matches the table shown at the
OpenHands Index Space.
from datasets import load_dataset
# leaderboard (one row per model)
ds = load_dataset("OpenHands/openhands-index", split="test")
ds.info.version # → "2026.06.30-3015ac6"
# per-instance outcomes (one row per model × instance)
instances = load_dataset("OpenHands/openhands-index"… See the full description on the dataset page: https://huggingface.co/datasets/OpenHands/openhands-index.SWE-Zero-openhands-trajectories
SWE-Zero Trajectories: Execution-free Fine-tuning for Software Engineering Agents
Data Overview
SWE-ZERO Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 318k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/SWE-Zero-openhands-trajectories.SWE-bench_Pro-locagentSWE-rebench-openhands-trajectories
Dataset Summary
SWE-rebench-OpenHands-Trajectories is a dataset of multi-turn agent trajectories for software engineering tasks, collected
using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.54.0) agent scaffolding.
This dataset captures complete agent execution traces as they attempt to resolve real GitHub issues from
nebius/SWE-rebench.
Each trajectory contains the agent's step-by-step reasoning, actions, and environmental observations.
Metric… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/SWE-rebench-openhands-trajectories.SWE-Hero-openhands-trajectories
SWE-Hero Trajectories: Execution-based Fine-tuning for Software Engineering Agents
Data Overview
SWE-Hero Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 34k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/SWE-Hero-openhands-trajectories.openhands-divergence-dpo-strong
openhands-divergence-dpo-strong
Strong-preference subset of divergence-point DPO pairs for OpenHands-style
tool use. Filtered for decisive exclusivity and chosen-strong actions
(edit/create/test/make_test/script; search only when the rejected side is
explore). Soft explore↔explore junk is out.
Current revision: strong-v2.2 (local soft post-filter after Round 12
spot-check fail). Remine quality bars unchanged from v1/v2. Volume from 8+8
re-spill; quality recovery drops… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/openhands-divergence-dpo-strong.swerebench-openhands-100m-max64k
SWE-rebench OpenHands 100M SFT Subset (max 64k)
This is a deterministic, representative subset of
nebius/SWE-rebench-openhands-trajectories,
augmented with exact sequence and supervised-loss token counts. The source trajectories were
collected with Qwen3-Coder-480B-A35B-Instruct and OpenHands v0.54.0. This derivative preserves
the source dataset's CC BY 4.0 license and attribution.
Filters and size
7,867 trajectories
100,095,655 assistant loss tokens
352,709,237… See the full description on the dataset page: https://huggingface.co/datasets/hanspeterlyngsoeraaschoujensen/swerebench-openhands-100m-max64k.swe-rebench-v2-qwen38-27b-bo3-openhands
SWE-Rebench-V2 Qwen3.8-27B BO3 annotations
This is a compact annotation overlay for nebius/SWE-rebench-V2 at revision
10c49ab856fd0e62097815ba5909dfc4f31e7f93. It summarizes campaign qwen38-swe-rebench-verified-all-bo3-ctx262k-seed20260822 using model openai/Qwen/Qwen3.8-27B.
It does not redistribute gold patches, test patches, raw trajectories, exception
text, credentials, service endpoints, or machine-local paths.
Incomplete snapshot — do not report this as final pass@3. Only… See the full description on the dataset page: https://huggingface.co/datasets/apurvaga/swe-rebench-v2-qwen38-27b-bo3-openhands.CodeScout_Training_Rolloutsopenhands_ce_data_v6terminal_bench_2_a1_swegym_openhands_20260808_162055
