Epoch-1
Datasets
All datasets matching “Epoch-1”details_Lansechen__Qwen2.5-3B-Instruct-Distill-om220k-1k-simplified-50b-batch32-epoch1-8192
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Instruct-Distill-om220k-1k-simplified-50b-batch32-epoch1-8192
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Instruct-Distill-om220k-1k-simplified-50b-batch32-epoch1-8192.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Instruct-Distill-om220k-1k-simplified-50b-batch32-epoch1-8192.acm-browsecompplus-train-rollouts-qwen3.5-9b-epoch1
lixiaochuan2020/acm-browsecompplus-train-rollouts-qwen3.5-9b-epoch1
BrowseComp-Plus train680 pass@4 rollouts (MemTool regime) from qwen3.5-9b-opd_iter1 — one row per (question, rep) in rollouts.jsonl.
Fields: question, gold, final_answer, correct_gpt5 (GPT-5 judge), num_turns, tool_call_counts, and full trajectories (raw_history + folded history + mem_operations + token_trajectory).
Stats: 2720 rollouts / 680 questions / reps [1, 2, 3, 4] · pass@1 70.0% · pass@4 83.2% (GPT-5).
details_Lansechen__Qwen2.5-3B-Instruct-Distill-om220k-2k-simplified-batch32-epoch1-8192
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Instruct-Distill-om220k-2k-simplified-batch32-epoch1-8192
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Instruct-Distill-om220k-2k-simplified-batch32-epoch1-8192.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 14 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Instruct-Distill-om220k-2k-simplified-batch32-epoch1-8192.20260722_3D_test_epoch1acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-1
Annotates the rollouts of: base model rollouts (react+memtool ×4)
One .npz per (question, rep) trajectory · 1145 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1.details_Lansechen__Qwen2.5-7B-Instruct-Distill-om220k-1k-origin-batch32-epoch1-16384
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Instruct-Distill-om220k-1k-origin-batch32-epoch1-16384
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Instruct-Distill-om220k-1k-origin-batch32-epoch1-16384.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Instruct-Distill-om220k-1k-origin-batch32-epoch1-16384.
