NiuNiu0110/RST-SFT-Qwen3.5-27B
RST SFT trajectories for Qwen3.5-27B Multi-turn terminal-agent conversations distilled from Zhongzhi1228/Recursive-Task-Synthesis-Trajectories, ready for supervised fine-tuning of Qwen/Qwen3.5-27B. Pipeline, launchers, and the full plan: https://github.com/k1ssloo/RST-Train cap10 reproduces the paper's SFT example count exactly The source release has 327,189 trajectories. cap10 ends at 10,778 examples — the count arXiv:2608.05466v3 states it trained on. That was… See the full description on the dataset page: https://huggingface.co/datasets/NiuNiu0110/RST-SFT-Qwen3.5-27B.
Correct the loss_mask note: do NOT shift (HF/Liger shift internally)
Upload README.md with huggingface_hub
Upload manifest_cap10_pretokenized.json with huggingface_hub
Upload data/cap10_pretokenized/holdout.parquet with huggingface_hub
Upload data/cap10_pretokenized/train.parquet with huggingface_hub
Upload README.md with huggingface_hub
Upload manifest_cap8.json with huggingface_hub
Upload data/cap8/holdout.parquet with huggingface_hub
Upload data/cap8/train.parquet with huggingface_hub
Upload manifest_cap10.json with huggingface_hub
Upload data/cap10/holdout.parquet with huggingface_hub
Upload data/cap10/train.parquet with huggingface_hub
initial commit
