CoolFace
Modelpublic

CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-OPD-iter149-0917

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes231downloads
Model Card

Qwen3-1.7B browsing worker — opd, iteration 149

Full BF16 evaluation export of qwen1p7b-sft1366-opd-teacher8b-historical-i149-0917/hf/iter_149_eval. This is the exact checkpoint used in the 2026-09-18 evaluation: all weight SHA-256 hashes match the frozen evaluation manifest. It derives from the SFT iteration 1366 worker and the user-provided opd training run. Training-method names are provenance; training hyperparameters and data overlap were not independently audited by the evaluation agent.

Base tokenizer: Qwen/Qwen3-1.7B at 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e. Context: 40,960 tokens. The evaluated export has 311 BF16 tensors. Padded embedding/head rows were trimmed during export; see eval-preparation.json and evaluation-provenance.json. The uploaded files are unchanged from that evaluated export, not a new conversion. SHA256SUMS covers source checkpoint files.

Evaluation

These are worker-under-orchestrator results, not solo-model accuracy. Concurrency is 64 top-level episodes. BrowseComp: 150 questions, 3 nodes, TP4 orchestrator and four independent TP1 workers. DR9K: 256 questions, 2 nodes, TP4 orchestrator and TP1/native DP4 workers. Full-answer repository manual grading protocol; not the canonical Gemini semantic judge. Empty/error episodes count wrong.

OrchestratorBrowseCompDR9K
MiniMax-M2.798/150 (65.33%)110/256 (42.97%)
DeepSeek-V4-Flash96/150 (64.00%)147/256 (57.42%)
Nemotron Ultra99/150 (66.00%)143/256 (55.86%)
MiMo-V2.5104/150 (69.33%)142/256 (55.47%)
Nemotron Super88/150 (58.67%)119/256 (46.48%)
Inkling-Small99/150 (66.00%)152/256 (59.38%)

Complete report includes latency settings, difficulty scores, observed role tokens/cache and separate theoretical worker prefix reuse.

BrowseComp traces · DR9K traces.

Use the report's evaluator, benchmark-specific chat templates, search tools and budgets to reproduce the worker results; generic chat generation alone does not reproduce this setup.