BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-RL-Math
Open-MOPD-SmolLM3-3B-RL-Math
This is the math-domain teacher in the Open-MOPD pipeline. It starts from BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-MixSFT and is trained only on math prompts with verifiable rewards using GRPO. This release corresponds to training step 100.
Training uses global batch size 128, mini-batch size 32, constant learning rate 1e-6 with 10 warmup steps, clipping at 0.2/0.25, rollout group size 16, temperature 1.0, a 30,000-token response limit, and no KL penalty. Groups with all-correct or all-incorrect generations are filtered, with up to eight resampling attempts.
Results
Math results use avg@64 with temperature 0.6. The broader evaluation setup uses max_model_len=32768, top_p=0.95, top_k=-1, and stop_token_ids=[128012].
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-RL-Math"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")Intended use and limitations
This is a domain teacher intended for distillation, not a general-purpose assistant. It was optimized only on math and can perform worse than MixSFT on other domains.
Model specifications
- Architecture:
SmolLM3ForCausalLM - Parameters: approximately 3B
- Layers: 36
- Vocabulary size: 128,256
- Weights: BF16, approximately 6.2 GB
- Includes tokenizer and chat template
