spreadsheet
SpreadsheetBench-v2
SpreadsheetBench 2
Home Page
SpreadsheetBench 2 is a benchmark for evaluating agents on end-to-end business spreadsheet workflows.
Unlike existing benchmarks that focus on isolated manipulations, SpreadsheetBench 2 requires agents to
(1) complete workflow-level goals through multi-step coordinated operations,
(2) perform cross-sheet reasoningwithin complex multi-sheet workbooks,
(3) produce deliverable-level outcomes including structured models, repaired spreadsheets, and… See the full description on the dataset page: https://huggingface.co/datasets/KAKA22/SpreadsheetBench-v2.SpreadsheetBench
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
| Paper | Github | Homepage |
We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed
to immerse current large language models (LLMs) in the actual workflow of spreadsheet users. Unlike existing benchmarks that rely on
synthesized queries and simplified spreadsheet files, SpreadsheetBench is built from 912 real questions gathered… See the full description on the dataset page: https://huggingface.co/datasets/KAKA22/SpreadsheetBench.Spreadsheet-RL
Spreadsheet-RL Dataset
Project Page | Paper | GitHub | Model
This dataset contains the training and evaluation data used by Spreadsheet-RL, a reinforcement learning framework for spreadsheet agents that edit Excel workbooks with tools and receive outcome-based rewards from workbook recalculation and answer-range comparison.
News
🚀 2026-08-01: Released the Spreadsheet-RL-8B checkpoint, scaling SpreadsheetBench Pass@1 from 15.9% for the base model to 16.7%… See the full description on the dataset page: https://huggingface.co/datasets/Spreadsheet-RL/Spreadsheet-RL.SpreadsheetBench-2-sample
SpreadsheetBench 2 Dataset
SpreadsheetBench 2 evaluates spreadsheet agents on end-to-end business and financial workflows. The release contains 321 tasks in four categories: Debugging, Financial Model, Template, and Visualization.
Overview
Directory
Task type
Tasks
Input workbooks
Gold workbooks
Debugging
Spreadsheet debugging and error correction
100
100
10
Financial_Model
Completion of multi-sheet financial models
100
100
20
Template
Completion of… See the full description on the dataset page: https://huggingface.co/datasets/icyCreater/SpreadsheetBench-2-sample.spreadsheet-arena-release
Spreadsheet Arena
A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models.
This is the public release accompanying the Spreadsheet Arena paper.
Contents
battles.csv
models.csv
outputs/<id>/
sheet.json
sheet.xlsx
<id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.spreadsheet-bench-v2-modified
SpreadsheetBench V2 Modified: Multi-Document QA
1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents.
An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.
