datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabular-logs-and-datasetstabular-data-wf9uh
Tabular Data Wf9Uh
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
3,251
Validation
409
Test
206
Total
3,866
Classes (12)
bold_parent_row
bold_row
closure_row
column
direct_children
non_bold_parent_row
non_bold_row
parent_column
prime_parent
sub_row
table
Usage
With LibreYOLO
from libreyolo… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/tabular-data-wf9uh.multitask-tabular-datasetsThis is a port of the Multi-Label Classification Dataset Repository (link).
We convert the datasets from there to simple csvs, resulting in 32 csvs (many of their mulan files fail to parse into python for us)
The targets in each csv are labeled with the suffix __target
Dataset
Domain
m
d
q
Card
Dens
Div
avgIR
rDep
m×q×d
0
3s-bbc1000
Text
352
1000
6
1.125
0.188
0.234
1.718
0.733
2.11e+06
1
3s-guardian1000
Text
302
1000
6
1.126
0.188
0.219
1.773
0.667
1.81e+06
2
3s-inter3000… See the full description on the dataset page: https://huggingface.co/datasets/imodels/multitask-tabular-datasets.IFT-Data-For-Tabular-Tasksyash-gym-tabular-dataset
Yash Gym Tabular Dataset
Dataset Summary
This dataset contains information on 30 unique gym machines with 5 consistent features and a binary target (Upper/Lower).It includes:
original: 30 manually collected samples
augmented: ~300 synthetic samples created with jitter, SMOTE-NC, MixUp, and CTGAN.
Intended Use
Educational dataset for tabular ML tasks, demonstrating preprocessing + augmentation.Not suitable for prescribing exercise or medical advice.… See the full description on the dataset page: https://huggingface.co/datasets/ysakhale/yash-gym-tabular-dataset.Tabular_DataSetsyalta_ai_tabular_dataset
YALTAi Tabular Dataset
353 page images of historical documents with tabular layouts, mostly notarial registers, with bounding-box annotations for four zone types: Col, Header, Marginal and text. Created by Thibault Clérice and deposited on Zenodo alongside the paper You Actually Look Twice At it (YALTAi) (Journal of Data Mining and Digital Humanities, 2022), which treats page layout recognition on historical documents as an object detection problem rather than a pixel… See the full description on the dataset page: https://huggingface.co/datasets/biglam/yalta_ai_tabular_dataset.hw1-tabular-hand-data
24-679 (Fall 2026): CMU Right-Hand Measurements
cmuchancel/hw1-tabular-hand-data
Right-hand finger measurements from Carnegie Mellon University students, plus explicitly marked
synthetic training variants. The classroom regression task predicts middle-finger length from thumb,
index-finger, ring-finger, and pinky lengths together with the recorded Female / Male categorical
feature. All finger measurements are stored in centimeters.
Source and task
The original… See the full description on the dataset page: https://huggingface.co/datasets/cmuchancel/hw1-tabular-hand-data.2026-24679-tabular-dataset
24-679 (Fall 2026): Music Listening Survey
ccm/2026-24679-tabular-dataset
Course-survey responses about music listening, plus explicitly marked synthetic training variants.
The classroom regression task predicts weekly listening hours from two count features and five
categorical preferences. Stored haiku answers provide source context and are excluded from this
task's predictors.
Source and task
The preparation notebook reads 24-679-tabular-survey.csv, removes the… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2026-24679-tabular-dataset.2026-24679-tabular-dataset
24-679 (Fall 2026): Snack Nutrition Dataset
kwongnon/2026-24679-tabular-dataset
Nutrition-label information for 50 packaged snack products. Each row represents one snack and
contains its product name, servings per container, calories, total fat, cholesterol, sodium,
total carbohydrate, and protein.
The dataset can be used for classroom exercises involving tabular data exploration, preprocessing,
visualization, clustering, regression, or other machine-learning tasks using… See the full description on the dataset page: https://huggingface.co/datasets/kwongnon/2026-24679-tabular-dataset.foods-tabular-datasetdataset_095129839_robotics_text_tabular
dataset_095129839_robotics_text_tabular.py
Dataset Summary
A robotics dataset with text tabular modality, stored in huggingface format.
Preprocessing & Augmentation
Preprocessing: adaptive
Augmentation: randaugment
Splits & Sampling
Split strategy: kfold 5
Sampling: stratified
Quality & Labeling
Quality filtering: strict
Labeling: pseudo label
Files
dataset_095129839_robotics_text_tabular.py — main… See the full description on the dataset page: https://huggingface.co/datasets/almutce90/dataset_095129839_robotics_text_tabular.test_composite_and_tabular_datasetdataset_095140933_robotics_text_tabular
dataset_095140933_robotics_text_tabular.py
Dataset Summary
A robotics dataset with text tabular modality, stored in csv format.
Preprocessing & Augmentation
Preprocessing: progressive
Augmentation: none
Splits & Sampling
Split strategy: stratified 90 10
Sampling: active
Quality & Labeling
Quality filtering: lenient
Labeling: self training
Files
dataset_095140933_robotics_text_tabular.py — main… See the full description on the dataset page: https://huggingface.co/datasets/JOONHOAHN/dataset_095140933_robotics_text_tabular.dataset_095058792_robotics_text_tabular
dataset_095058792_robotics_text_tabular.py
Dataset Summary
A robotics dataset with text tabular modality, stored in hdf5 format.
Preprocessing & Augmentation
Preprocessing: auto ml
Augmentation: heavy
Splits & Sampling
Split strategy: temporal
Sampling: contrastive
Quality & Labeling
Quality filtering: moderate
Labeling: weak supervision
Files
dataset_095058792_robotics_text_tabular.py — main… See the full description on the dataset page: https://huggingface.co/datasets/abdullahalotaibi/dataset_095058792_robotics_text_tabular.text-tabular-dataset
Sound Events Text Tabular Data Notes
Dataset summary
This data card accompanies a lightweight Sound Events loader for Text Tabular metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/donghyunkwon34k/text-tabular-dataset.dataset_095133318_robotics_text_tabular
dataset_095133318_robotics_text_tabular.py
Dataset Summary
A robotics dataset with text tabular modality, stored in webdataset format.
Preprocessing & Augmentation
Preprocessing: adaptive
Augmentation: randaugment
Splits & Sampling
Split strategy: temporal
Sampling: curriculum
Quality & Labeling
Quality filtering: lenient
Labeling: self training
Files
dataset_095133318_robotics_text_tabular.py —… See the full description on the dataset page: https://huggingface.co/datasets/michaelmmm/dataset_095133318_robotics_text_tabular.dataset_141139128_math_text_tabular
dataset_141139128_math_text_tabular.py
Dataset Summary
A math dataset with text tabular modality, stored in arrow format.
Preprocessing & Augmentation
Preprocessing: auto ml
Augmentation: heavy
Splits & Sampling
Split strategy: random 90 10
Sampling: active
Quality & Labeling
Quality filtering: moderate
Labeling: self training
Files
dataset_141139128_math_text_tabular.py — main artifact of this… See the full description on the dataset page: https://huggingface.co/datasets/daniilvasil/dataset_141139128_math_text_tabular.dataset_018187735_ecommerce_text_tabular
dataset_018187735_ecommerce_text_tabular.py
Dataset Summary
A ecommerce dataset with text tabular modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: minimal
Augmentation: randaugment
Splits & Sampling
Split strategy: kfold 5
Sampling: active
Quality & Labeling
Quality filtering: moderate
Labeling: manual
Files
dataset_018187735_ecommerce_text_tabular.py — main… See the full description on the dataset page: https://huggingface.co/datasets/williamlpux83/dataset_018187735_ecommerce_text_tabular.dataset_002963818_medical_text_tabular
dataset_002963818_medical_text_tabular.py
Dataset Summary
A medical dataset with text tabular modality, stored in huggingface format.
Preprocessing & Augmentation
Preprocessing: adaptive
Augmentation: autoaugment
Splits & Sampling
Split strategy: temporal
Sampling: random
Quality & Labeling
Quality filtering: moderate
Labeling: self training
Files
dataset_002963818_medical_text_tabular.py — main… See the full description on the dataset page: https://huggingface.co/datasets/sakuraksk95/dataset_002963818_medical_text_tabular.text-tabular-dataset
Education Text Tabular Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Education work with Text Tabular inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/bdegroot1982/text-tabular-dataset.dataset_095099426_robotics_text_tabular
dataset_095099426_robotics_text_tabular.py
Dataset Summary
A robotics dataset with text tabular modality, stored in lmdb format.
Preprocessing & Augmentation
Preprocessing: adaptive
Augmentation: light
Splits & Sampling
Split strategy: random 90 10
Sampling: stratified
Quality & Labeling
Quality filtering: moderate
Labeling: semi auto
Files
dataset_095099426_robotics_text_tabular.py — main artifact… See the full description on the dataset page: https://huggingface.co/datasets/jjohnsonner/dataset_095099426_robotics_text_tabular.dataset_052867206_traffic_text_tabular
dataset_052867206_traffic_text_tabular.py
Dataset Summary
A traffic dataset with text tabular modality, stored in webdataset format.
Preprocessing & Augmentation
Preprocessing: adaptive
Augmentation: heavy
Splits & Sampling
Split strategy: stratified 90 10
Sampling: random
Quality & Labeling
Quality filtering: moderate
Labeling: semi auto
Files
dataset_052867206_traffic_text_tabular.py — main… See the full description on the dataset page: https://huggingface.co/datasets/felixlehmann/dataset_052867206_traffic_text_tabular.dataset_048842448_sports_text_tabular
dataset_048842448_sports_text_tabular.py
Dataset Summary
A sports dataset with text tabular modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: mixup cutmix
Splits & Sampling
Split strategy: leave one out
Sampling: stratified
Quality & Labeling
Quality filtering: moderate
Labeling: self training
Files
dataset_048842448_sports_text_tabular.py —… See the full description on the dataset page: https://huggingface.co/datasets/tiffanysanchez/dataset_048842448_sports_text_tabular.dataset_052663079_traffic_text_tabular
dataset_052663079_traffic_text_tabular.py
Dataset Summary
A traffic dataset with text tabular modality, stored in arrow format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: none
Splits & Sampling
Split strategy: random 90 10
Sampling: weighted
Quality & Labeling
Quality filtering: adaptive
Labeling: pseudo label
Files
dataset_052663079_traffic_text_tabular.py — main artifact… See the full description on the dataset page: https://huggingface.co/datasets/haozhangley/dataset_052663079_traffic_text_tabular.database-small-tabular-regressiondataset_141096971_math_text_tabular
dataset_141096971_math_text_tabular.py
Dataset Summary
A math dataset with text tabular modality, stored in csv format.
Preprocessing & Augmentation
Preprocessing: domain specific
Augmentation: mixup cutmix
Splits & Sampling
Split strategy: leave one out
Sampling: hard negative
Quality & Labeling
Quality filtering: strict
Labeling: semi auto
Files
dataset_141096971_math_text_tabular.py — main… See the full description on the dataset page: https://huggingface.co/datasets/trevorhallton/dataset_141096971_math_text_tabular.hw1-24679-tabular-dataset
Shoe Size Measurements Tabular Dataset
Dataset Summary
Purpose: This dataset was created for tabular data analysis and prediction tasks involving shoe measurements, developed as part of CMU 24-679 coursework to explore tabular data augmentation techniques.
Quick Stats:
338 total samples (30 original + 308 augmented)
3 numerical features + 3 categorical features
High correlation between size measurements (>0.97)
~10x augmentation factor
Contact: maryzhang@cmu.edu… See the full description on the dataset page: https://huggingface.co/datasets/maryzhang/hw1-24679-tabular-dataset.dataset_079869779_education_text_tabular
dataset_079869779_education_text_tabular.py
Dataset Summary
A education dataset with text tabular modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: curriculum
Augmentation: randaugment
Splits & Sampling
Split strategy: random 90 10
Sampling: contrastive
Quality & Labeling
Quality filtering: lenient
Labeling: pseudo label
Files… See the full description on the dataset page: https://huggingface.co/datasets/amyhughes/dataset_079869779_education_text_tabular.
