datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tabmwptabm-dataDatasets used in the paper: TabM: Advancing Tabular Deep Learning With Parameter-Efficient Ensembling
Download data:
wget https://huggingface.co/datasets/rototoHF/tabm-data/resolve/main/data.tar
Tab-MIA
Tab-MIA: A Benchmark for Membership Inference Attacks on Tabular Data
Tab-MIA is a benchmark dataset designed to evaluate the privacy risks of fine-tuning large language models (LLMs) on structured tabular data. It enables reproducible and systematic testing of Membership Inference Attacks (MIAs) across diverse datasets and six different serialization formats.
📋 Overview
Datasets:
WTQ (WikiTableQuestions)
WikiSQL
TabFact
Adult Census
California Housing… See the full description on the dataset page: https://huggingface.co/datasets/germane/Tab-MIA.tabmwp-cleanTabMI-Bench
TabMI-Bench
A protocol benchmark for mechanistic interpretability (MI) of tabular foundation models (TFMs). NeurIPS 2026 Evaluations & Datasets Track submission.
What's in this dataset
This Hugging Face repository hosts the frozen aggregated artifacts that drive every numbered table and figure in the paper. Bundling these allows reviewers to verify the paper's key numerics without re-running 40 GPU-hours of experiments.
File
Source experiment
Used by… See the full description on the dataset page: https://huggingface.co/datasets/EvalData/TabMI-Bench.tabmwptabmwp_expel_train_100tabmwp-hard-verifiedfineweb-sample-10BT-completion-augmented-v0all-the-newss2orc-academic-papers-augmented-v0tabmwp_200tabmwp_200_trainfineweb-sample-10BT-augmented-v0all_the_news_augmented-v0tabmaven-270925finepdfs-eng_Latn-augmented-v0
