datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RST-SFT-Qwen3.5-27B
RST SFT trajectories for Qwen3.5-27B
Multi-turn terminal-agent conversations distilled from
Zhongzhi1228/Recursive-Task-Synthesis-Trajectories,
ready for supervised fine-tuning of Qwen/Qwen3.5-27B.
Pipeline, launchers, and the full plan: https://github.com/k1ssloo/RST-Train
cap10 reproduces the paper's SFT example count exactly
The source release has 327,189 trajectories. cap10 ends at 10,778 examples —
the count arXiv:2608.05466v3 states it trained
on. That was… See the full description on the dataset page: https://huggingface.co/datasets/NiuNiu0110/RST-SFT-Qwen3.5-27B.RST-DPO-Qwen3.5-27B
RST DPO preference pairs for Qwen3.5
2,673 preference pairs (2,448 train / 225 holdout) built from
Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.
Each pair is two agent runs on the same task: one whose trajectory the task's own
verifier scored reward 1, one it scored 0.
Builder, trainer, and the numerical gates: https://github.com/k1ssloo/RST-Train
(scripts/17_build_dpo_data.py, scripts/19_train_dpo.py, DPO_PLAN.md).
What the preference actually encodes… See the full description on the dataset page: https://huggingface.co/datasets/NiuNiu0110/RST-DPO-Qwen3.5-27B.
