datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabMWPtable_spill_cleanup_bimanual
Exylos Bimanual Spill Cleanup — Rich-Modality 50-Episode Sample
50 episodes of a bimanual Franka Panda wiping a liquid spill off a tabletop. Synthetic, VR-teleop demonstrations retargeted to two 7-DoF arms — 6 RGB views (3 with depth + segmentation), 6-DoF object poses, and a ground-truth dirty_fraction cleanliness signal, packaged in LeRobot v2.1.
Release note: this rich-modality v2 release replaces the original public 50-episode preview in place. The previous dataset… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual.WikiTableQuestionsevaluation-tables
[!CAUTION]
This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025.
elements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
bird-dev-tablesArXiv-tables
Arxiv-tables Dataset
Dataset Summary
The Arxiv-tables dataset is a collection of tables extracted from scientific papers published on arXiv, primarily focused on ML papers. It includes both the LaTeX source of the tables and their corresponding rendered images from the PDF versions of the papers.
Supported Tasks
This dataset can support several tasks, including but not limited to:
Table structure recognition
LaTeX to image generation for tables
Image-to-LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/staghado/ArXiv-tables.tpch_tables_scale_1
polars-tpch
This repo contains the code used for performance evaluation of polars. The benchmarks are TPC-standardised queries and data designed to test the performance of "real" workflows.
From the TPC website:
TPC-H is a decision support benchmark. It consists of a suite of business-oriented ad hoc queries and concurrent data modifications. The queries and the data populating the database have been chosen to have broad industry-wide relevance. This benchmark illustrates… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/tpch_tables_scale_1.helpmate-tables
Helpmate tablebases
Exhaustively solved helpmate
tablebases: for every legal position in a material class, the distance to
mate under optimal cooperation and the number of distinct optimal
solutions.
That second number is the useful one. It is what tells a composer whether a
position is a sound problem (count = 1) or has duals.
Generated by
osick/helpmate-tablebase
(MIT). These files are the data; that repository is the code that builds and
reads them.
629 six-piece tablebases… See the full description on the dataset page: https://huggingface.co/datasets/osick/helpmate-tables.SciMolmo_images_tables_2025TABLET-tables
TABLET-tables
Here you can find the image and HTML files for all tables in the TABLET dataset.These tables are already included in all TABLET datasets so you don't need to download anything in this repository if you only want to use the datasets.TABLET datasets on Hugging Face are self-contained (each example already includes its table image and HTML), this resource provides access to the complete collection of table images and HTML files in one place.These files are… See the full description on the dataset page: https://huggingface.co/datasets/alonsoapp/TABLET-tables.tables-3.4
[!NOTE]
Dataset origin: http://infolingu.univ-mlv.fr/
Description
Lexicon-Grammar tables of French verbs, nouns playing the predicative role ("verbal nouns"), frozen expressions and adverbs in CSV format.For more details about the modifications of tables, see (E. Tolone, 2009), (E. Tolone, 2011) and (E. Tolone, 2012).The archive file contains:Simple verbs:
67 tables with 13 872 entries, including 5 738 distinct entries
thae table of classes with 552 features
arag index of… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/tables-3.4.FreeformTableQATableSuite-1K
TableSuite-1K
TableSuite-1K benchmarks predictive and language-grounded tabular intelligence
over 1,000 OpenML-referenced datasets.
Task
Input
Evaluation
Prediction
ICL rows or a partially labelled serialized table
classification and regression
Table grounding
a provided table plus a lookup/comprehension question
exact displayed-table facts
Table QA
a provided subtable plus a typed question
programmatic operations
This repository contains metadata and… See the full description on the dataset page: https://huggingface.co/datasets/Lester1996/TableSuite-1K.table_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.AgiBotWorld-Beta_G1_task_470_Organize_the_tables_in_the_milk_tea_shop
agibot_task_470
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 整理奶茶店的桌子。
total_episodes: 661
total_tasks: 1
size: 47G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_470_Organize_the_tables_in_the_milk_tea_shop.paperswithcode-data-evaluation-tables
Process data from paperswithcode
See https://huggingface.co/datasets/pwc-archive/files/tree/main.
Download and unzip evaluation tables:
curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz"
gunzip jul-28-evaluation-tables.json.gz
Install jq.
See https://jqlang.org/.
If on Debian/Ubuntu, install with sudo apt-get install jq.
Example jq to extract:
jq -r '
def process(parent):
.task as $current_task |
(if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.AgiBotWorld-Beta_G1_task_378_Clean_up_the_tables_in_the_restaurant
agibot_task_378
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 清理餐厅的桌子
total_episodes: 739
total_tasks: 1
size: 47G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_378_Clean_up_the_tables_in_the_restaurant.WikiTableQuestionsSelectionThis dataset is a curated subset of the original WikiTableQuestions (Pasupat and Liang, 2015). To address inconsistencies and inaccuracies found in the source material—such as ambiguous queries and incorrect ground-truth labels—this version consists of 100 hand-selected examples. Each entry has been verified to ensure high data quality and factual alignment, making it an ideal benchmark for precise table-based QA evaluation.
open_tables_icttd_for_table_detectionDatasets for the paper "Revisiting Table Detection Datasets for Visually Rich Documents" (https://arxiv.org/abs/2305.04833) (https://link.springer.com/article/10.1007/s10032-025-00527-9).
Benchmark
We buidt a new benchmark with this dataset. Please refer to https://github.com/uobinxiao/SparseTableDet for the details.
License
Since this dataset is built on several open datasets and open documents, users should also adhere to the licenses of these publicly available… See the full description on the dataset page: https://huggingface.co/datasets/uobinxiao/open_tables_icttd_for_table_detection.table_smoke
V10.10 table-camera validation set
Ten pick-and-place demonstrations from the pact_place_corridor_v10_10_four_object environment, re-run with one
extra fixed exterior table camera (exo_camera_1) so the camera schema can be
validated end to end before it is used at scale.
This is a schema validation set, not a training corpus. The 144-row V10.10
training corpus is unchanged and is not part of this repo.
What each row guarantees
Every published row was checked… See the full description on the dataset page: https://huggingface.co/datasets/Lundii/table_smoke.TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.twfilter-tables
twfilter reference tables
The tables that decide whether a span of traditional-Chinese text is Taiwanese Mandarin*
(臺灣華語, cmn-Hant-TW) rather than Hong Kong Cantonese, mainland text converted to traditional
characters, or literary Chinese. Plain text, one record per line, tab-separated where a
record has fields, LC_ALL=C sort order, UTF-8, LF.
Consumed by twfilter 0.1.0, where this
directory is vendored byte-for-byte and MANIFEST.json is verified by its test suite.
Usable… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twfilter-tables.tables-pdf-20260823-pagespubmed-tables-latex-768pxbrick-skill-tables
Brick public skill tables
Public skill vectors consumed by the Brick router. The Hugging Face dataset
regolo/brick-skill-tables contains one CSV file, skill_vectors.csv, with one
row per model and six capability values in [0,1]. Brick uses these values as
cold-start priors, so users do not need to measure a model that is already listed.
The CLI also ships richer JSON copies under this folder for offline initialization.
The Hugging Face dataset is intentionally CSV-only;… See the full description on the dataset page: https://huggingface.co/datasets/regolo/brick-skill-tables.pretrain_tables_mergedipo-tables
IPO Tables (HTML) — Random Sample Card
A curated table extraction dataset from SEC filing documents, with raw table HTML plus provenance metadata.
What This Dataset Is
This is a random sample targeting 100 extracted tables per year from filings in 1994–2026.
Middle years are densely represented at 100 tables/year.
Edge years can be lower where fewer valid tables were available.
Tables are extracted directly from filing source files and stored as raw HTML.
Full… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-tables.table_scViRL39K-Tables-Diagrams-Charts
