CoolFace
Modelpublic

HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-003

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes572downloads
Model Card

OnlineRubrics RaR-Medicine: step 3, seed 11

Intermediate policy from dynamic OnlineRubrics-Every GRPO training. Distinct from static-rubric GRPO. Base model: Qwen/Qwen3-4B-Instruct-2507; thinking disabled. This checkpoint is a policy state used by the Phase-1 audit. No downstream medical capability or safety claim is made. Research use only; not validated for clinical decision-making.

Root files are the veRL-exported Hugging Face inference model (BF16). original_checkpoint/ preserves the exact original FSDP parameter checkpoint and tokenizer/configuration files. Optimizer state, training data, responses, rubrics, infrastructure configuration, and credentials are not included. The original is retained because export precision/serialization differs.

Base model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Original actor tree SHA256: c2dd930f30c54f9b962b0cc17669871a89d1e71f2abcb748cf38fe115a53eddb