CoolFace
12 results

officeqa

databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K26 likes9k downloads2mo agoHugging FaceNEAR-AI /officeqa OfficeQA manifest (nearai-bench packaging) Harness-ready question manifest for databricks/officeqa — document-grounded QA over U.S. Treasury Bulletins (1939–2025). 246 items in full, 8 in smoke (a 4-easy/4-hard subset for pipeline checks). from datasets import load_dataset ds = load_dataset("NEAR-AI/officeqa", split="full") ⚠️ This is the manifest only — documents are NOT included Unlike our pinchbench and clawbench exports, the source corpus is not bundled here.… See the full description on the dataset page: https://huggingface.co/datasets/NEAR-AI/officeqa.textquestion-answeringn<1K0 likes6.9k downloads2mo agoHugging Facedatabricks /officeqa-pro-v2gated OfficeQA Pro v2 Dataset Summary OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.documentquestion-answeringn<1K17 likes2.5k downloads2mo agoHugging FaceYWZBrandon /officeqa-checkpoint-eval-data Checkpoint evaluation plot data Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures. No model execution, grading, publication, or source-result changes were performed to make this export. Contents checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds. pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.tabular10K<n<100K0 likes1.5k downloads8d agoHugging Face