datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabMWPWikiTableQuestionsReasoning-Table
Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
The Reasoning-Table dataset is a high-quality, reasoning dataset designed for table reasoning tasks.
📁 Directory Structure
This repository is organized by task. Each subfolder contains task-specific reasoning data, including raw and filtered versions. Here is an overview:
├── fetaqa/
├── feverous/
├── finqa/
├── gsm8k/
├── hitab/
├── hybridqa/
├── multihierttt/
├── ottqa/
├── tabfact/
├── tatqa/
├──… See the full description on the dataset page: https://huggingface.co/datasets/TableQAKit/Reasoning-Table.Table-GPT
Table-GPT: Table-tuned GPT for Diverse Table Tasks
This repository contains training and test datasets for the SIGMOD'24 paper Table-GPT: Table-tuned GPT for Diverse Table Tasks. The source code for data generation and task evaluation are available here: https://github.com/microsoft/Table-GPT, which can be used to generate more training data for table-related tasks.
Task Descriptions
We collect (or synthesize) 18 diverse table-related tasks, which are summarized in… See the full description on the dataset page: https://huggingface.co/datasets/LipengCS/Table-GPT.2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value while the traces are ~3x longer, by keeping only a subset of the expanded rows. Its sibling arm holds the ROW share instead; together they separate more deliberation from more difficult-advice signal.
field
value
experiment
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture.FreeformTableQAWTQTableInstruct
Citation
@misc{wu2024tablebenchcomprehensivecomplexbenchmark,
title={TableBench: A Comprehensive and Complex Benchmark for Table Question Answering},
author={Xianjie Wu and Jian Yang and Linzheng Chai and Ge Zhang and Jiaheng Liu and Xinrun Du and Di Liang and Daixin Shu and Xianfu Cheng and Tianzhen Sun and Guanglin Niu and Tongliang Li and Zhoujun Li},
year={2024},
eprint={2408.09174},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/TableInstruct.WikiTableQuestionsSelectionThis dataset is a curated subset of the original WikiTableQuestions (Pasupat and Liang, 2015). To address inconsistencies and inaccuracies found in the source material—such as ambiguous queries and incorrect ground-truth labels—this version consists of 100 hand-selected examples. Each entry has been verified to ensure high data quality and factual alignment, making it an ideal benchmark for precise table-based QA evaluation.
parsebench-table-track
ParseBench Table Track plus Financial Split
This dataset is a mirror of the table dimension of llamaindex/ParseBench, packaged together with a curated Financial Split that we built for evaluating OCR systems on insurance and financial filings.
It contains,
All 503 PDFs of the ParseBench table track.
table.jsonl, the original ground truth (one HTML table per page, plus easy or hard difficulty tag).
financial_split/, our 151 page financial slice plus the 117 dropped non financial… See the full description on the dataset page: https://huggingface.co/datasets/roma2025/parsebench-table-track.TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.SQA2026-08-26-table2-9284-low-stakes-716-train
Low-stakes difficult advice, mixed for training (10,000)
The swap-in twin of LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train. Same 9,284
benign rows, same 716 scenarios, same renderer, same seed — the difficult-advice half
replaced by its low-stakes rewrite. Train this against that control and the only thing
that differs is the magnitude of what the 716 scenarios put at risk.
field
value
experiment
Low-stakes arm of difficult advice: does lowering the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-table2-9284-low-stakes-716-train.html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.2026-08-04-table2-synthdoc-h200x4-train
Training bundle: Table2 (8,000) + synthdoc (2,202), 4xH200 DDP
code.tar.gz (trainer, src/, configs/) plus mixture_think.jsonl (10,202 rows).
Non-reasoning rows carry the empty think marker as CONTEXT; mask_empty_think: true
excludes those tokens from the loss. synthdoc rows keep real reasoning traces and are
supervised. Config: configs/train/2026-08-25_lora_qwen36_table2_synthdoc.yaml.
2026-08-27-table2-9284-good-ai-fiction-716-train
Table2 9,284 + Good AI Fiction 716 — SFT training mixture
field
value
experiment
The fiction arm of the alignment-data comparison: the SAME 9,284 benign capability-preserving rows the difficult-advice mixture uses, with its 716 difficult-advice rows replaced by 716 first-person Good AI Fiction rows at a matched trainable-token budget. Train against LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train to read the difference as content, not size.… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-table2-9284-good-ai-fiction-716-train.2026-08-04-table2-only-9284-h200x4-train
Training bundle: Table 2 only, 9,284 examples (no difficult-advice)
code.tar.gz plus mixture_think.jsonl. Every row carries the empty think marker as
context; mask_empty_think: true keeps those tokens out of the loss.
Config: configs/train/2026-08-25_lora_qwen36_table2_only.yaml.
hitab2026-08-17-table2-9284-peer-critique-good-716-train-mixture
Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows)
The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose
716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm.
field
value
experiment
Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm.
date_generated
2026-08-17
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.2026-08-25-table2-9284-difficult-advice-verbose-716-train
Verbose-CoT arm: the published difficult-advice mixture with the 716 difficult-advice reasoning traces expanded ~3x in length, same ideas, to isolate deliberation length from content.
field
value
experiment
Verbose-CoT arm: the published difficult-advice mixture with the 716 difficult-advice reasoning traces expanded ~3x in length, same ideas, to isolate deliberation length from content.
date_generated
20260825
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-716-train.2026-09-02-table2-9284-nonmoral-deliberation-684-train-mixture
Training mixture for the DELIBERATION-WITHOUT-MORALITY arm: 9,284 spec-filtered Table-2 instruction rows + 684 rows in which an assistant is handed a concrete piece of work with a binary instruction about how to do it, the specifics make that instruction the worse call, and the assistant says so and does it its way. Nothing moral is at stake in any of the 684: nobody is harmed, deceived, endangered or treated unfairly. The swap-in twin of… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-02-table2-9284-nonmoral-deliberation-684-train-mixture.2026-08-06-gpqa-diamond-table2-onlyFreeformTableQASelectionThis dataset represents a curated subset of the FreeformTableQA collection, refined to provide a more rigorous benchmark for table-based reasoning. While the original dataset covers a broad range of "free-form" queries—which often include complex, semi-structured, or non-grid layouts—it also contains instances of noise and misaligned labels. To ensure higher evaluation accuracy, this version features 100 manually verified examples where the natural language queries and tabular evidence have… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/FreeformTableQASelection.2026-08-28-table2-9284-par716coh-train
Training mixture for the COHERENT post-action-retrospection 716 arm (arm 1 of the PAR coherence experiment): the same 9,284 spec-filtered Table-2 rows as every sibling arm + the SAME 716 five-turn PAR rows that trained the par716 arm, with only their trained turn rewritten so the reasoning ends on a first-person decision and the reply enacts it (LASR-Callum/2026-08-28-post-action-retrospection-716-coherent). Row-for-row paired with… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-table2-9284-par716coh-train.Rotowire-Text-to-Table
RotoWire Corrected Test Set
Dataset Description
This is the corrected test set for the RotoWire dataset, released as part of the Map&Make: Schema Guided Text to Table Generation paper (ACL 2025). The original RotoWire dataset contained hallucination errors where the ground truth tables included incorrect or fabricated statistics. This corrected version provides a cleaner benchmark for text-to-table generation tasks.
Motivation: Why We Need This Corrected… See the full description on the dataset page: https://huggingface.co/datasets/McH04/Rotowire-Text-to-Table.punditbench-2026-27-season-tables
PunditBench 2026-27 pre-registered season tables
How did 40–42 language models rank every club in Europe's five largest football leagues before the season began?
This static dataset contains 204 complete final-table forecasts from PunditBench's 2026–27 league benchmark. Each row is one model's predicted finishing order for one league. Every included forecast was locked before that league's opening kickoff, committed publicly, hashed, and tagged in Git.
La Liga: 40 tables… See the full description on the dataset page: https://huggingface.co/datasets/skebbe/punditbench-2026-27-season-tables.2026-09-01-table2-9284-difficult-advice-rewritten-702-train-mixture
Table2-9284 + REWRITTEN difficult-advice-702 (thinking SFT mixture)
Rewritten variant of the principle-scoped difficult-advice-702 mixture: the 9,284 standard SFT rows are
byte-identical; the 702 difficult-advice rows have their REASONING and ANSWER rewritten.
field
value
experiment
1-epoch LoRA SFT mixture for Qwen3.6-27B: 9,284 Table2 rows + 702 rewritten difficult-advice rows (7.03% synthetic). Rewrite (on top of the earlier ablation): reasoning OPENING -> neutral… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-table2-9284-difficult-advice-rewritten-702-train-mixture.TableTex
TableTex
Dataset Summary
TableTex is a dataset for table image-to-LaTeX generation, comprising
19,000 renderable LaTeX table annotations extracted from scientific articles
on arXiv. Each annotation corresponds to a table image that can be generated
locally using the provided generate_images.py script. The rendered images
are not included in the repository, which keeps the distributed dataset
compact while providing a reproducible image-generation pipeline.
To… See the full description on the dataset page: https://huggingface.co/datasets/yunfanyang1/TableTex.snowfox-financial-table-data
snowfox-financial-table-data
SnowFox — financial-table extraction training (snowfox_financial_table tasks).
Contents
train.jsonl (1680 rows)
validation.jsonl (162 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for SnowFox (Michael Anthony Falabella).
US_Native_American_Tribal_Treaties_Table_from_WikipediaORIGINALLY COMPILED ON WIKIPEDIA. Cleaned and improved from original version on Wikipedia, by completing some column information that was available elsewhere in the wikipedia article.
This is a dataset of tabular data regarding treaties between the USA and Native American Tribes/Nations, to date, including many executive orders.
Table columns include Year, Date, Treaty name, "Alternative Treaty name", Statutes, "Land cession reference (Royce Area)", Tribe(s).
All of those listed include the… See the full description on the dataset page: https://huggingface.co/datasets/pseudolab/US_Native_American_Tribal_Treaties_Table_from_Wikipedia.
