CoolFace
Modelpublic

budget-internalization-iclr2027/qwen3.5-4b-8k-dapo-bosstroll-s300

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes13downloads
Model Card

Qwen3.5-4B · 8k-token budget · dapo · bosstroll

RL-finetuned Qwen/Qwen3.5-4B trained for math reasoning under a 8,192-token generation budget. Released as part of an anonymous ICLR 2027 submission.

Run ID (petname): `bosstroll` · checkpoint step 300

Training

Base modelQwen/Qwen3.5-4B
AlgorithmDAPO: clip-higher (clip range 0.2 / 0.28), dynamic sampling (prompt groups with identical rewards are dropped and resampled), token-level loss, soft overlong penalty (up to -1) over the last 1,024 tokens before the budget, rewards scaled to [-1, 1], no leave-one-out baseline.
Generation budget (max_new_tokens)8,192
DataDeepScaleR (math), 3 epochs max
Batch32 prompts × 8 rollouts per step
OptimizerAdam, cosine LR schedule, peak LR 5e-7, 10 warmup steps
Steps300
Rewardbinary answer correctness (\boxed{} extraction)
Weights dtypeBF16

Training prompt (user turn, rendered with the base model's chat template):

Think step-by-step to solve the following problem. Output your answer inside of \\boxed{} tags.:
{problem}

Let's think step-by-step

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "budget-internalization-iclr2027/qwen3.5-4b-8k-dapo-bosstroll-s300"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")

With vLLM: vllm serve budget-internalization-iclr2027/qwen3.5-4b-8k-dapo-bosstroll-s300

License

Inherits the license of the base model (Qwen/Qwen3.5-4B).