CoolFace
Modelpublic

hyun1905/qwen3-4b-instruct-2507-review-point-oneshot-sft

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes1.2kdownloads
Model Card

Qwen3-4B One-Shot Review-Point SFT

This checkpoint is the initial policy for one-shot review-point GRPO. It was supervised-fine-tuned to read a paper and emit up to nine high-quality review points as one numbered list, with Qwen thinking disabled.

The expected user prompt has this shape:

text
# Paper
<paper text>

# Instruction
Write a list of high-quality review points for the paper.
...

The assistant output is a numbered list with one review point per item.

This is a research checkpoint. Its review points and factual claims require independent verification before consequential use.