CoolFace
Modelpublic

Spreadsheet-RL/Spreadsheet-RL-4B

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
7likes60downloads
Model Card

Spreadsheet-RL-4B

<p align="center"> <img src="spreadsheet-rl.png" alt="Spreadsheet-RL logo" width="700"> </p>

**Project Page** | **Paper** | **Dataset** | **Code**

Spreadsheet-RL-4B is the RL-trained 4B spreadsheet agent checkpoint from Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning. It starts from `Qwen/Qwen3-4B-Thinking-2507` and is post-trained with outcome-based reinforcement learning in Spreadsheet Gym, a multi-turn Microsoft Excel environment with spreadsheet-native tools, sandboxed code execution, and Excel-based recalculation rewards.

This checkpoint is intended to be used with the Spreadsheet-RL agent harness and tool environment. Loading it as a plain chat model can be useful for inspection, but it will not reproduce the paper results without Spreadsheet Gym, the tool set, and the reward/evaluation pipeline.

News

Model Details

FieldValue
Base model`Qwen/Qwen3-4B-Thinking-2507`
Training methodGRPO with outcome-based rewards
EnvironmentSpreadsheet Gym with Microsoft Excel 365, spreadsheet-native tools, SandboxFusion code execution, and async Excel recalculation/reward service
Training dataSpreadsheet-RL training split: 5,928 filtered ExcelForum tasks
EvaluationSpreadsheetBench and Domain-Spreadsheet
LicenseApache-2.0, following the base model license

Training Configuration

For full details, please see the paper. The released 4B run uses:

HyperparameterValue
AlgorithmGRPO; KL-regularized against a frozen reference model
Training steps60
Prompt/response limits4,096 / 27,648 tokens
Rollout samplingtemperature 0.6; top-p 0.95; top-k 20
Batching64 prompts/step; 16 rollouts/prompt; 1,024 rollouts/step
Multi-turn capsmax assistant turns 20; max user turns 20; max tool-response length 8,192
OptimizerAdamW; learning rate 1e-6; weight decay 0.01; betas (0.9, 0.999); grad clip 1.0
KL losslow-var KL; coefficient 0.001
Actor update batchingmini-batch 32; dynamic batch sizing enabled
Hardware1 node x 4 NVIDIA H100 GPUs
Training timeabout 40 hours wall-clock for the 4B run

Results

Spreadsheet-RL improves the same 4B base model through spreadsheet-native interaction design, comprehensive tool access, and RL post-training.

BenchmarkBase+ Native Harness+ Full ToolsSpreadsheet-RL-4B
SpreadsheetBench Pass@112.015.619.323.4

On Domain-Spreadsheet, Spreadsheet-RL improves overall Pass@1 from 8.4 to 17.2 over 1,660 evaluation rollouts.

Usage

Install the standard Transformers stack and load the checkpoint:

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Spreadsheet-RL/Spreadsheet-RL-4B"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)

For task evaluation and agent rollouts, use the full Spreadsheet-RL codebase with the released dataset and Spreadsheet Gym:

bash
hf download Spreadsheet-RL/Spreadsheet-RL --repo-type dataset --local-dir data
git clone https://github.com/Spreadsheet-RL/Spreadsheet-RL.git

The default training/evaluation harness is maintained in the code repository under configs/, scripts/, reward/, and verl/.

Citation

bibtex
@misc{chi2026spreadsheetrl,
  title         = {Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning},
  author        = {Banghao Chi and Yining Xie and Mingyuan Wu and Jingcheng Yang and Jize Jiang and Zhaoheng Li and Shengyi Qian and Minjia Zhang and Klara Nahrstedt and Rui Hou and Xiangjun Fan and Hanchao Yu},
  year          = {2026},
  eprint        = {2605.22642},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  doi           = {10.48550/arXiv.2605.22642},
  url           = {https://arxiv.org/abs/2605.22642}
}