CoolFace
Datasetpublic

kilian-group/KBevo-SFT-hotpotqa-6k

KBevo-SFT-hotpotqa-6k Supervised fine-tuning (SFT) trajectories for the KBevo two-phase policy, generated on HotpotQA. Accompanies Co-Evolving Structured Knowledge and Reasoning in Language Models (COLM 2026). This is the exact SFT dataset used to produce kilian-group/KBevo-Qwen3-1.7B-SFT and kilian-group/KBevo-Qwen3-4B-SFT, which in turn initialise the KBevo-Qwen3-1.7B-GRPO and KBevo-Qwen3-4B-GRPO runs. What's in the file A single JSON file, trajectories.json… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/KBevo-SFT-hotpotqa-6k.

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
0likes52downloads
4 commits on main
22d8cc714d ago

Fix card: annotated_text is the SFT target (not full_response); 11746 records over 5744 unique questions; each record is a full two-phase trajectory

zhaolx19
7ddc38514d ago

Upload SFT trajectories (11746 rollouts)

zhaolx19
8bd85d314d ago

Add dataset card

zhaolx19
0582f9d14d ago

initial commit

zhaolx19