CoolFace
Modelpublic

CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-Base-GRPO-iter149-0916

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes230downloads
Model Card

Qwen3-1.7B browsing worker — base-grpo, iteration 149

Full BF16 evaluation export of qwen1p7b-sft1366-base-grpo-g8b8-lr1e6-g36flash-0916/hf/iter_149_eval. This is the exact checkpoint used in the 2026-09-18 evaluation: all weight SHA-256 hashes match the frozen evaluation manifest. It derives from the SFT iteration 1366 worker and the user-provided base-grpo training run. Training-method names are provenance; training hyperparameters and data overlap were not independently audited by the evaluation agent.

Base tokenizer: Qwen/Qwen3-1.7B at 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e. Context: 40,960 tokens. The evaluated export has 311 BF16 tensors. Padded embedding/head rows were trimmed during export; see eval-preparation.json and evaluation-provenance.json. The uploaded files are unchanged from that evaluated export, not a new conversion. SHA256SUMS covers source checkpoint files.

Evaluation

These are worker-under-orchestrator results, not solo-model accuracy. Concurrency is 64 top-level episodes. BrowseComp: 150 questions, 3 nodes, TP4 orchestrator and four independent TP1 workers. DR9K: 256 questions, 2 nodes, TP4 orchestrator and TP1/native DP4 workers. Full-answer repository manual grading protocol; not the canonical Gemini semantic judge. Empty/error episodes count wrong.

OrchestratorBrowseCompDR9K
MiniMax-M2.7100/150 (66.67%)120/256 (46.88%)
DeepSeek-V4-Flash99/150 (66.00%)148/256 (57.81%)
Nemotron Ultra98/150 (65.33%)143/256 (55.86%)
MiMo-V2.5113/150 (75.33%)153/256 (59.77%)
Nemotron Super92/150 (61.33%)128/256 (50.00%)
Inkling-Small110/150 (73.33%)158/256 (61.72%)

Complete report includes latency settings, difficulty scores, observed role tokens/cache and separate theoretical worker prefix reuse.

BrowseComp traces · DR9K traces.

Use the report's evaluator, benchmark-specific chat templates, search tools and budgets to reproduce the worker results; generic chat generation alone does not reproduce this setup.