kilian-group/KBevo-SFT-hotpotqa-6k
KBevo-SFT-hotpotqa-6k Supervised fine-tuning (SFT) trajectories for the KBevo two-phase policy, generated on HotpotQA. Accompanies Co-Evolving Structured Knowledge and Reasoning in Language Models (COLM 2026). This is the exact SFT dataset used to produce kilian-group/KBevo-Qwen3-1.7B-SFT and kilian-group/KBevo-Qwen3-4B-SFT, which in turn initialise the KBevo-Qwen3-1.7B-GRPO and KBevo-Qwen3-4B-GRPO runs. What's in the file A single JSON file, trajectories.json… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/KBevo-SFT-hotpotqa-6k.
Fix card: annotated_text is the SFT target (not full_response); 11746 records over 5744 unique questions; each record is a full two-phase trajectory
Upload SFT trajectories (11746 rollouts)
Add dataset card
initial commit
