officeqa
qwen3.5-9b-officeqa-pure-opd-last-it139qwen3.5-9b-officeqa-aod2op-tail-last-it139qwen3.5-4b-officeqa-pure-opd-best-it59qwen3.5-4b-officeqa-pure-opd-last-it139qwen3.5-4b-officeqa-aod2op-tail-last-it139qwen3.5-9b-officeqa-pure-opd-best-it79qwen3.5-9b-officeqa-aod2op-tail-best-it90qwen3.5-27b-officeqa-pure-opd-best-it49
Datasets
All datasets matching “officeqa”officeqa
OfficeQA
Dataset Summary
OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.officeqa
OfficeQA manifest (nearai-bench packaging)
Harness-ready question manifest for
databricks/officeqa — document-grounded
QA over U.S. Treasury Bulletins (1939–2025). 246 items in full,
8 in smoke (a 4-easy/4-hard subset for pipeline checks).
from datasets import load_dataset
ds = load_dataset("NEAR-AI/officeqa", split="full")
⚠️ This is the manifest only — documents are NOT included
Unlike our pinchbench and
clawbench exports, the
source corpus is not bundled here.… See the full description on the dataset page: https://huggingface.co/datasets/NEAR-AI/officeqa.officeqa-pro-v2
OfficeQA Pro v2
Dataset Summary
OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.officeqa-checkpoint-eval-data
Checkpoint evaluation plot data
Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures.
No model execution, grading, publication, or source-result changes were performed to make this export.
Contents
checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds.
pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.
