datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabMWPWikiTableQuestionsFreeformTableQAWikiTableQuestionsSelectionThis dataset is a curated subset of the original WikiTableQuestions (Pasupat and Liang, 2015). To address inconsistencies and inaccuracies found in the source material—such as ambiguous queries and incorrect ground-truth labels—this version consists of 100 hand-selected examples. Each entry has been verified to ensure high data quality and factual alignment, making it an ideal benchmark for precise table-based QA evaluation.
TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.FreeformTableQASelectionThis dataset represents a curated subset of the FreeformTableQA collection, refined to provide a more rigorous benchmark for table-based reasoning. While the original dataset covers a broad range of "free-form" queries—which often include complex, semi-structured, or non-grid layouts—it also contains instances of noise and misaligned labels. To ensure higher evaluation accuracy, this version features 100 manually verified examples where the natural language queries and tabular evidence have… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/FreeformTableQASelection.punditbench-2026-27-season-tables
PunditBench 2026-27 pre-registered season tables
How did 40–42 language models rank every club in Europe's five largest football leagues before the season began?
This static dataset contains 204 complete final-table forecasts from PunditBench's 2026–27 league benchmark. Each row is one model's predicted finishing order for one league. Every included forecast was locked before that league's opening kickoff, committed publicly, hashed, and tagged in Git.
La Liga: 40 tables… See the full description on the dataset page: https://huggingface.co/datasets/skebbe/punditbench-2026-27-season-tables.prompts_for_tables_normalization_and_new_sqlsmd-2-xml-wiki-tables
md-2-xml-wiki-tables
958 markdown tables extracted from fan/community MediaWiki sites for markdown-to-XML format conversion tasks.
Format
JSONL with fields:
title: article title from the source wiki page
section: section heading the table appeared under
wiki: source wiki name
table_md: raw markdown table
filename: original filename
Splits
train: 894 tables
eval: 64 held-out tables
Source
Various fan/community MediaWiki sites. Most use CC-BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/md-2-xml-wiki-tables.table-sft-eval-predictions
💾 Raw Predictions for "What Really Matters for Table LLMs?"
This dataset contains the raw model outputs from the experiments in:
Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang,
Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng.
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects.
Findings of EACL 2026. https://aclanthology.org/2026.findings-eacl.195/
🗂️ Layout… See the full description on the dataset page: https://huggingface.co/datasets/dnaihao/table-sft-eval-predictions.hecm-plf-tables
HECM Principal Limit Factor (PLF) Tables
Complete PLF lookup table with 4,864 age-rate combinations for calculating reverse mortgage borrowing capacity. Ages 62-99, interest rates 3.0%-18.875%.
How PLF Works
PLF determines what percentage of your home's value you can access:
Higher age = higher PLF (more money available)
Lower interest rate = higher PLF
Example: Age 72 at 6.5% rate = PLF of 0.434 (43.4% of home value)
Data Fields
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/hecm-plf-tables.simple-between-tables-clariden-conversions
SIMPLE BetweenTables Clariden conversions
This dataset is the dedicated publication boundary for SIMPLE→SONIC conversions produced and accepted on CSCS Clariden for G1WholebodyLocomotionPickBetweenTablesTeleop-v0.
Current status
Reserved; no accepted generation is published yet.
The submitted Clariden pilot must complete and pass fresh source-free replay before any token payload is uploaded. Existing lab-machine campaigns from minimax, monotone, or optimality are… See the full description on the dataset page: https://huggingface.co/datasets/dlsmarta/simple-between-tables-clariden-conversions.relationalrag-tables-artifact-data
RelationalRAG-Tables Artifact Data
This public dataset repository mirrors the lightweight data payloads from the
RelationalRAG-Tables reproducibility artifact:
fixtures/: smoke and synthetic fixtures used by the CPU-only artifact path.
artifacts/: stored JSON outputs and manifest.json hashes for reported
experimental numbers.
The repository does not redistribute heavyweight upstream benchmark corpora
such as BIRD or HybridQA, nor does it redistribute pretrained model weights.… See the full description on the dataset page: https://huggingface.co/datasets/lexuanbach/relationalrag-tables-artifact-data.f_up_tablesExplainMed-lookup-tables
