datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabMWPtabmwpTABMEpp
Dataset Card for TABME++
The TABME dataset is a synthetic collection of business document folders generated from the Truth Tobacco Industry Documents archive, with preprocessing and OCR results included, designed to simulate real-world digitization tasks.
TABME++ extends TABME by enriching it with commercial-quality OCR (Microsoft OCR).
Dataset Details
Dataset Description
The TABME dataset is a synthetic collection created to simulate the… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/TABMEpp.TabMWPtabm-dataDatasets used in the paper: TabM: Advancing Tabular Deep Learning With Parameter-Efficient Ensembling
Download data:
wget https://huggingface.co/datasets/rototoHF/tabm-data/resolve/main/data.tar
tabmwparxiv-rawTab-MIA
Tab-MIA: A Benchmark for Membership Inference Attacks on Tabular Data
Tab-MIA is a benchmark dataset designed to evaluate the privacy risks of fine-tuning large language models (LLMs) on structured tabular data. It enables reproducible and systematic testing of Membership Inference Attacks (MIAs) across diverse datasets and six different serialization formats.
📋 Overview
Datasets:
WTQ (WikiTableQuestions)
WikiSQL
TabFact
Adult Census
California Housing… See the full description on the dataset page: https://huggingface.co/datasets/germane/Tab-MIA.TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.s2orc-academic-papers-rawtabmwp-cleanTabMWPmm_tabmwpTabMI-Bench
TabMI-Bench
A protocol benchmark for mechanistic interpretability (MI) of tabular foundation models (TFMs). NeurIPS 2026 Evaluations & Datasets Track submission.
What's in this dataset
This Hugging Face repository hosts the frozen aggregated artifacts that drive every numbered table and figure in the paper. Bundling these allows reviewers to verify the paper's key numerics without re-running 40 GPU-hours of experiments.
File
Source experiment
Used by… See the full description on the dataset page: https://huggingface.co/datasets/EvalData/TabMI-Bench.TabMCQtabmwpVQA-tabmwp
Description
This dataset is a processed version of the TabMWP dataset by Lu et al.We converted the images to PIL and translated question and answers from Englist to French.
Citation
TabMWP
@misc{lu2023dynamicpromptlearningpolicy,
title={Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning},
author={Pan Lu and Liang Qiu and Kai-Wei Chang and Ying Nian Wu and Song-Chun Zhu and Tanmay Rajpurohit and Peter Clark and… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/VQA-tabmwp.tabmwp_expel_train_100tabmwp_200fineweb-sample-10BT-completion-augmented-v0tabmwp_cleaned
tabmwp_cleaned
The tabmwp__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
24,966
QA turns
143,078
answers rewritten by the cleaning pass
6,128
QA created by the cleaning pass (new_qa)
93,315 (65.2%)
shards
1
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/tabmwp_cleaned.all-the-newss2orc-academic-papers-augmented-v0tabmwp-hard-verifiedtabmwp_200fineweb-sample-10BT-augmented-v0tabmwp_200_trainall_the_news_augmented-v0tabmaven-270925tabme_small
