budget-internalization-iclr2027/qwen3.5-4b-4k-grpo-budget-system-viableroughy-s250
015
Qwen3.5-4B · 4k-token budget · grpo-budget-system · viableroughy
RL-finetuned Qwen/Qwen3.5-4B trained for math reasoning under a 4,096-token generation budget that is stated to the model in a system prompt. Released as part of an anonymous ICLR 2027 submission.
Run ID (petname): `viableroughy` · checkpoint step 250
Note: this run was stopped at step 250 of the planned 300; other checkpoints in the budget sweep are at step 300.
Training
Prompts are DeepScaleR problems with a statement of the 4,096-token budget placed in a system prompt, rendered with the base model's chat template.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "budget-internalization-iclr2027/qwen3.5-4b-4k-grpo-budget-system-viableroughy-s250"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")With vLLM: vllm serve budget-internalization-iclr2027/qwen3.5-4b-4k-grpo-budget-system-viableroughy-s250
License
Inherits the license of the base model (Qwen/Qwen3.5-4B).
