budget-internalization-iclr2027/qwen3.5-4b-2k-grpo-lean-upwardgull-s50
012
Qwen3.5-4B · 2k-token budget · grpo-lean · upwardgull
RL-finetuned Qwen/Qwen3.5-4B trained with GRPO on Lean Workbook formal theorem proving under a 2,048-token generation budget. Released as part of an anonymous ICLR 2027 submission.
Run ID (petname): `upwardgull` · checkpoint step 50
Training
Prompts are the Lean Workbook formal theorem proving task prompts, rendered with the base model's chat template.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "budget-internalization-iclr2027/qwen3.5-4b-2k-grpo-lean-upwardgull-s50"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")With vLLM: vllm serve budget-internalization-iclr2027/qwen3.5-4b-2k-grpo-lean-upwardgull-s50
License
Inherits the license of the base model (Qwen/Qwen3.5-4B).
