datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabMWPWikiTableQuestionstable_spill_cleanup_bimanual
Exylos Bimanual Spill Cleanup — Rich-Modality 50-Episode Sample
50 episodes of a bimanual Franka Panda wiping a liquid spill off a tabletop. Synthetic, VR-teleop demonstrations retargeted to two 7-DoF arms — 6 RGB views (3 with depth + segmentation), 6-DoF object poses, and a ground-truth dirty_fraction cleanliness signal, packaged in LeRobot v2.1.
Release note: this rich-modality v2 release replaces the original public 50-episode preview in place. The previous dataset… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual.evaluation-tables
[!CAUTION]
This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025.
elements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
bird-dev-tablesArXiv-tables
Arxiv-tables Dataset
Dataset Summary
The Arxiv-tables dataset is a collection of tables extracted from scientific papers published on arXiv, primarily focused on ML papers. It includes both the LaTeX source of the tables and their corresponding rendered images from the PDF versions of the papers.
Supported Tasks
This dataset can support several tasks, including but not limited to:
Table structure recognition
LaTeX to image generation for tables
Image-to-LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/staghado/ArXiv-tables.tpch_tables_scale_1
polars-tpch
This repo contains the code used for performance evaluation of polars. The benchmarks are TPC-standardised queries and data designed to test the performance of "real" workflows.
From the TPC website:
TPC-H is a decision support benchmark. It consists of a suite of business-oriented ad hoc queries and concurrent data modifications. The queries and the data populating the database have been chosen to have broad industry-wide relevance. This benchmark illustrates… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/tpch_tables_scale_1.FreeformTableQATableSuite-1K
TableSuite-1K
TableSuite-1K benchmarks predictive and language-grounded tabular intelligence
over 1,000 OpenML-referenced datasets.
Task
Input
Evaluation
Prediction
ICL rows or a partially labelled serialized table
classification and regression
Table grounding
a provided table plus a lookup/comprehension question
exact displayed-table facts
Table QA
a provided subtable plus a typed question
programmatic operations
This repository contains metadata and… See the full description on the dataset page: https://huggingface.co/datasets/Lester1996/TableSuite-1K.table_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.paperswithcode-data-evaluation-tables
Process data from paperswithcode
See https://huggingface.co/datasets/pwc-archive/files/tree/main.
Download and unzip evaluation tables:
curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz"
gunzip jul-28-evaluation-tables.json.gz
Install jq.
See https://jqlang.org/.
If on Debian/Ubuntu, install with sudo apt-get install jq.
Example jq to extract:
jq -r '
def process(parent):
.task as $current_task |
(if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.WikiTableQuestionsSelectionThis dataset is a curated subset of the original WikiTableQuestions (Pasupat and Liang, 2015). To address inconsistencies and inaccuracies found in the source material—such as ambiguous queries and incorrect ground-truth labels—this version consists of 100 hand-selected examples. Each entry has been verified to ensure high data quality and factual alignment, making it an ideal benchmark for precise table-based QA evaluation.
brick-skill-tables
Brick public skill tables
Public skill vectors consumed by the Brick router. The Hugging Face dataset
regolo/brick-skill-tables contains one CSV file, skill_vectors.csv, with one
row per model and six capability values in [0,1]. Brick uses these values as
cold-start priors, so users do not need to measure a model that is already listed.
The CLI also ships richer JSON copies under this folder for offline initialization.
The Hugging Face dataset is intentionally CSV-only;… See the full description on the dataset page: https://huggingface.co/datasets/regolo/brick-skill-tables.TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.twfilter-tables
twfilter reference tables
The tables that decide whether a span of traditional-Chinese text is Taiwanese Mandarin*
(臺灣華語, cmn-Hant-TW) rather than Hong Kong Cantonese, mainland text converted to traditional
characters, or literary Chinese. Plain text, one record per line, tab-separated where a
record has fields, LC_ALL=C sort order, UTF-8, LF.
Consumed by twfilter 0.1.0, where this
directory is vendored byte-for-byte and MANIFEST.json is verified by its test suite.
Usable… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twfilter-tables.pubmed-tables-latex-768pxpretrain_tables_mergedopenai-hf-incident-recovered-tables
Recovered tables from the METR OpenAI and Hugging Face incident figures
This is an unofficial third party dataset. METR did not produce it, review it or endorse it.
METR's report on the OpenAI and Hugging Face incident contains two interactive charts. Those
charts load their numbers from JavaScript files on metr.org. This dataset holds those numbers
as CSV tables, with the source URLs, the SHA-256 of each source file, and a script that
downloads the sources again and checks the… See the full description on the dataset page: https://huggingface.co/datasets/LauraGomezjurado/openai-hf-incident-recovered-tables.FreeformTableQASelectionThis dataset represents a curated subset of the FreeformTableQA collection, refined to provide a more rigorous benchmark for table-based reasoning. While the original dataset covers a broad range of "free-form" queries—which often include complex, semi-structured, or non-grid layouts—it also contains instances of noise and misaligned labels. To ensure higher evaluation accuracy, this version features 100 manually verified examples where the natural language queries and tabular evidence have… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/FreeformTableQASelection.table_scViRL39K-Tables-Diagrams-Chartsarocrbench_tablesPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
SEC-10Q-10K-Statement-tablesfintabnet-hq-tablesband-conditionality-tables
Band-Conditionality of a Reservoir Capacity Ranking
Derived result tables, run logs and figures for the methodological note Rank Order Under a
Capacity Benchmark Is Conditional on the Observation Band and the Delay Horizon
(Zharnikov 2026, concept DOI 10.5281/zenodo.22206844).
The configs: block in this card's frontmatter is load-bearing, not decoration. Each table
below has a different schema, so Hugging Face's default heuristics try to concatenate them into one
split and the… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/band-conditionality-tables.punditbench-2026-27-season-tables
PunditBench 2026-27 pre-registered season tables
How did 40–42 language models rank every club in Europe's five largest football leagues before the season began?
This static dataset contains 204 complete final-table forecasts from PunditBench's 2026–27 league benchmark. Each row is one model's predicted finishing order for one league. Every included forecast was locked before that league's opening kickoff, committed publicly, hashed, and tagged in Git.
La Liga: 40 tables… See the full description on the dataset page: https://huggingface.co/datasets/skebbe/punditbench-2026-27-season-tables.chartqa-tables
ChartQA Tables
This dataset contains pre-extracted tables and metadata from the ChartQA dataset by Ahmed Masry et al.
Dataset Description
ChartQA is a benchmark for question answering about charts with visual and logical reasoning. This companion dataset provides:
Structured tables extracted from chart images (CSV format)
Formatted tables in the paper's format for model input
Purpose
The original ChartQA paper evaluated models in two modes:
With gold tables… See the full description on the dataset page: https://huggingface.co/datasets/nmayorga7/chartqa-tables.SEC_Tables_Lite
Dataset Card for "SEC_Tables_Lite"
More Information needed
handbook-chemistry-physics-tables
