datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openhands-divergence-dpo-strong
openhands-divergence-dpo-strong
Strong-preference subset of divergence-point DPO pairs for OpenHands-style
tool use. Filtered for decisive exclusivity and chosen-strong actions
(edit/create/test/make_test/script; search only when the rejected side is
explore). Soft explore↔explore junk is out.
Current revision: strong-v2.2 (local soft post-filter after Round 12
spot-check fail). Remine quality bars unchanged from v1/v2. Volume from 8+8
re-spill; quality recovery drops… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/openhands-divergence-dpo-strong.openhands-divergence-dpo
openhands-divergence-dpo
5,420 divergence-point DPO pairs mined from
nebius/SWE-rebench-openhands-trajectories
(67,074 OpenHands v0.54.0 trajectories by Qwen3-Coder-480B-A35B-Instruct on real
GitHub issues from SWE-rebench).
Where Nebius released raw trajectories for RFT/RL, this dataset extracts
preference pairs at the first diverging action: for GitHub issues attempted
multiple times where at least one attempt resolved the issue and at least one
failed, we align a resolved and… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/openhands-divergence-dpo.2026.RA.Divergence-DPO-Pairs
Rational-Agent Divergence DPO Pairs
This is the corrected p4_dataset_v2 preference dataset built from the
validity-gated P2 negotiation campaign. Each row pairs the model's stored
per-turn rendered prompt and verbatim action with an executable
best-response-oracle action rendered in the same wire format.
All 12,761 rows contain the required experiment-name field, set to
p4-divergence-dpo-v2.
Splits
Split
Rows
Purpose
train
5,533
Episode-level training… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Divergence-DPO-Pairs.
