shanyangmie/physics-r1-seed23-canonical-step60-fsdp
Physics-R1 — Seed 23, canonical step-60 (FSDP-sharded)
**Project Page** | **Paper** | **Code** | **Training corpus**
The Physics-R1 paper checkpoint for the seed-23 row of Table 2. Fine-tune of Qwen3-VL-8B-Thinking on the audited `PhysR1Corp` (2,268 closed-form physics problems) via full-parameter FSDP1 GRPO with binary correctness reward.
Released alongside Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning.
Table 2 performance (this checkpoint, seed-23 row of paper Table 2)
Scoring: problem-level liberal Sonnet-as-judge (problem-AND aggregation; every subpart must be correct). See paper Appendix on judges.
The 3-seed mean across {42, 17, 23} is the paper's headline number — see Table 2 of the paper for full multi-seed mean ± σ.
Variants
Training recipe
- Base model: `Qwen/Qwen3-VL-8B-Thinking`
- Algorithm: GRPO (verl 0.6.1, full-parameter FSDP1 —
actor.strategy=fsdp, notfsdp2; FSDP2 fails on Qwen3-VL visual encoder device placement) - Reward: binary correctness, per-subpart Sonnet judge with problem-level AND aggregation (see paper §3.2)
- Data: `shanyangmie/physr1corp` — 2,268 audited closed-form problems
- Hardware: 4×H200 (FSDP1 4-way sharded)
Full hyperparameters in the paper appendix.
- Seed / step: 23 / 60
Format: verl FSDP-sharded checkpoint (conversion required)
This checkpoint is saved in verl's FSDP-sharded format, not safetensors. It is not directly loadable via AutoModelForImageTextToText.from_pretrained without a merge step.
File layout
actor/
├── huggingface/ # HF-style config + tokenizer
├── model_world_size_4_rank_{0,1,2,3}.pt # 4-way FSDP weight shards (~8.7 GB each, ~35 GB total)
├── optim_world_size_4_rank_{0,1,2,3}.pt # optimizer state (~17.5 GB each, not needed for inference)
├── extra_state_world_size_4_rank_{0..3}.pt
└── fsdp_config.json
data.pt # verl bookkeeping (not needed for inference)Convert to HF safetensors
Use verl's model_merger.py:
git clone https://github.com/volcengine/verl
cd verl
# Download only the inference-required files (skips ~70 GB of optimizer state)
huggingface-cli download shanyangmie/physics-r1-seed23-canonical-step60-fsdp \\
--include "actor/model_world_size_4_rank_*.pt" \\
--include "actor/huggingface/*" \\
--include "actor/fsdp_config.json" \\
--include "actor/extra_state_world_size_4_rank_*.pt" \\
--local-dir ./ckpt
# Merge FSDP shards into HF safetensors
python scripts/model_merger.py merge \\
--backend fsdp \\
--hf_model_path Qwen/Qwen3-VL-8B-Thinking \\
--local_dir ./ckpt/actor \\
--target_dir ./physics-r1-seed23-canonical-step60-fsdp-hfThen load with standard HF:
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(
"./physics-r1-seed23-canonical-step60-fsdp-hf",
torch_dtype="bfloat16",
device_map="auto",
)
processor = AutoProcessor.from_pretrained("./physics-r1-seed23-canonical-step60-fsdp-hf")License
Apache 2.0, inheriting from the base model `Qwen3-VL-8B-Thinking`. Training data (physr1corp) is CC BY-NC 4.0, so this derivative checkpoint is intended for non-commercial research use.
Citation
@misc{yang2026physicsr1,
title = {Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning},
author = {Yang, Shan},
year = {2026},
url = {https://huggingface.co/papers/2605.14040}
}