CoolFace
Modelpublic

ShuhaoChen202401/deepsearch-paper-filter-rl

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes11downloads
Model Card

DeepSearch Paper Filter — RL

Qwen3-8B checkpoint optimized with GRPO for continuous paper suitability scoring. It is released for the paper-filter training-strategy ablation; the default system uses the SFT paper filter.