CoolFace
Datasetpublic

lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3

lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3 Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime). Trains: OPD iter-3 Annotates the rollouts of: iter-2 rollouts (…-train-rollouts-…-epoch2) One .npz per (question, rep) trajectory · 736 files. Schema (per file, numpy.load) key shape dtype meaning input_ids (L,) int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes427downloads
Dataset Card

lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3

Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).

  • —Trains: OPD iter-3
  • —Annotates the rollouts of: iter-2 rollouts (…-train-rollouts-…-epoch2)
  • —One .npz per (question, rep) trajectory · 736 files.

Schema (per file, numpy.load)

keyshapedtypemeaning
input_ids(L,)int32full tokenized trajectory (student tokenizer Qwen/Qwen3.5-9B)
asst_pos(M,)int32positions where KD loss applies (assistant tokens)
tk_ids(M, 20)int32teacher's top-20 token ids at each of those positions
tk_logprobs(M, 20)float16teacher log-probs for those top-20 tokens

Filenames: <question_id>_run<rep>.npz.

Load

python
import numpy as np
z = np.load("1003_run1.npz")
input_ids, asst_pos, tk_ids, tk_lp = z["input_ids"], z["asst_pos"], z["tk_ids"], z["tk_logprobs"]
# KD: at each asst position, forward-KL between the renormalized teacher top-20 (tk_lp)
# and the student's log-probs gathered at tk_ids. See distill/train_kd.py + distill/kd_loss.py.