CoolFace
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asaverren /openhands-divergence-dpo-strong openhands-divergence-dpo-strong Strong-preference subset of divergence-point DPO pairs for OpenHands-style tool use. Filtered for decisive exclusivity and chosen-strong actions (edit/create/test/make_test/script; search only when the rejected side is explore). Soft explore↔explore junk is out. Current revision: strong-v2.2 (local soft post-filter after Round 12 spot-check fail). Remine quality bars unchanged from v1/v2. Volume from 8+8 re-spill; quality recovery drops… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/openhands-divergence-dpo-strong.texttext-generationn<1K0 likes91 downloads7d agoHugging Face02asaverren /openhands-divergence-dpo openhands-divergence-dpo 5,420 divergence-point DPO pairs mined from nebius/SWE-rebench-openhands-trajectories (67,074 OpenHands v0.54.0 trajectories by Qwen3-Coder-480B-A35B-Instruct on real GitHub issues from SWE-rebench). Where Nebius released raw trajectories for RFT/RL, this dataset extracts preference pairs at the first diverging action: for GitHub issues attempted multiple times where at least one attempt resolved the issue and at least one failed, we align a resolved and… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/openhands-divergence-dpo.tabulartext-generation1K<n<10K0 likes56 downloads11d agoHugging Face03siddharthmb /2026.RA.Divergence-DPO-Pairs Rational-Agent Divergence DPO Pairs This is the corrected p4_dataset_v2 preference dataset built from the validity-gated P2 negotiation campaign. Each row pairs the model's stored per-turn rendered prompt and verbatim action with an executable best-response-oracle action rendered in the same wire format. All 12,761 rows contain the required experiment-name field, set to p4-divergence-dpo-v2. Splits Split Rows Purpose train 5,533 Episode-level training… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Divergence-DPO-Pairs.text-generation10K<n<100K0 likes29 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.