ddvd233/medxpert9b_specgap_ship_retrieval_websearch_global_step_180
032
medxpert9bspecgapshipretrievalwebsearch (global step 180)
Merged HF weights (bf16 safetensors) of the verl FSDP checkpoint medxpert9b_specgap_ship_retrieval_websearch/global_step_180.
- Base model: Qwen/Qwen3.5-9B
- Experiment: RRIMed (self-evolving rewards), held-out arm on MedXpertQA (text + multimodal), ARM 21
- Score: MedXpertQA validation accuracy 0.407 (step 170 eval; run best 0.427 at step 70, whose weights were rotated away; untrained 0.357)
- Note: RL-only run from the base model with retrieval + web search tools; the reward is a one-criterion rubric so the grader yields exact-match accuracy.
This is the best checkpoint of the run whose weights survived checkpoint rotation; the run's best-validation step is stated above when it differs.
Research artifact trained on model-written tasks and evaluated on one benchmark family. Not for clinical use.
