CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-OPD-iter149-0917
Qwen3-1.7B browsing worker — opd, iteration 149
Full BF16 evaluation export of qwen1p7b-sft1366-opd-teacher8b-historical-i149-0917/hf/iter_149_eval. This is the exact checkpoint used in the 2026-09-18 evaluation: all weight SHA-256 hashes match the frozen evaluation manifest. It derives from the SFT iteration 1366 worker and the user-provided opd training run. Training-method names are provenance; training hyperparameters and data overlap were not independently audited by the evaluation agent.
Base tokenizer: Qwen/Qwen3-1.7B at 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e. Context: 40,960 tokens. The evaluated export has 311 BF16 tensors. Padded embedding/head rows were trimmed during export; see eval-preparation.json and evaluation-provenance.json. The uploaded files are unchanged from that evaluated export, not a new conversion. SHA256SUMS covers source checkpoint files.
Evaluation
These are worker-under-orchestrator results, not solo-model accuracy. Concurrency is 64 top-level episodes. BrowseComp: 150 questions, 3 nodes, TP4 orchestrator and four independent TP1 workers. DR9K: 256 questions, 2 nodes, TP4 orchestrator and TP1/native DP4 workers. Full-answer repository manual grading protocol; not the canonical Gemini semantic judge. Empty/error episodes count wrong.
Complete report includes latency settings, difficulty scores, observed role tokens/cache and separate theoretical worker prefix reuse.
BrowseComp traces · DR9K traces.
Use the report's evaluator, benchmark-specific chat templates, search tools and budgets to reproduce the worker results; generic chat generation alone does not reproduce this setup.
