budget-internalization-iclr2027/gemma4-e2b-2k-grpo-reasoninggym-smartcamel-s300
011
Gemma-4-E2B-it · 2k-token budget · grpo-reasoninggym · smartcamel
RL-finetuned google/gemma-4-E2B-it trained with GRPO on Reasoning Gym tasks under a 2,048-token generation budget. Released as part of an anonymous ICLR 2027 submission.
Run ID (petname): `smartcamel` · checkpoint step 300
Training
Prompts are the Reasoning Gym task prompts, rendered with the base model's chat template.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "budget-internalization-iclr2027/gemma4-e2b-2k-grpo-reasoninggym-smartcamel-s300"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")With vLLM: vllm serve budget-internalization-iclr2027/gemma4-e2b-2k-grpo-reasoninggym-smartcamel-s300
License
Inherits the license of the base model (google/gemma-4-E2B-it).
