lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3 Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime). Trains: OPD iter-3 Annotates the rollouts of: iter-2 rollouts (…-train-rollouts-…-epoch2) One .npz per (question, rep) trajectory · 736 files. Schema (per file, numpy.load) key shape dtype meaning input_ids (L,) int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3.
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
- Trains: OPD iter-3
- Annotates the rollouts of: iter-2 rollouts (…-train-rollouts-…-epoch2)
- One
.npzper (question, rep) trajectory · 736 files.
Schema (per file, numpy.load)
Filenames: <question_id>_run<rep>.npz.
Load
import numpy as np
z = np.load("1003_run1.npz")
input_ids, asst_pos, tk_ids, tk_lp = z["input_ids"], z["asst_pos"], z["tk_ids"], z["tk_logprobs"]
# KD: at each asst position, forward-KL between the renormalized teacher top-20 (tk_lp)
# and the student's log-probs gathered at tk_ids. See distill/train_kd.py + distill/kd_loss.py.