datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.yash-gym-tabular-dataset
Yash Gym Tabular Dataset
Dataset Summary
This dataset contains information on 30 unique gym machines with 5 consistent features and a binary target (Upper/Lower).It includes:
original: 30 manually collected samples
augmented: ~300 synthetic samples created with jitter, SMOTE-NC, MixUp, and CTGAN.
Intended Use
Educational dataset for tabular ML tasks, demonstrating preprocessing + augmentation.Not suitable for prescribing exercise or medical advice.… See the full description on the dataset page: https://huggingface.co/datasets/ysakhale/yash-gym-tabular-dataset.yalta_ai_tabular_dataset
YALTAi Tabular Dataset
353 page images of historical documents with tabular layouts, mostly notarial registers, with bounding-box annotations for four zone types: Col, Header, Marginal and text. Created by Thibault Clérice and deposited on Zenodo alongside the paper You Actually Look Twice At it (YALTAi) (Journal of Data Mining and Digital Humanities, 2022), which treats page layout recognition on historical documents as an object detection problem rather than a pixel… See the full description on the dataset page: https://huggingface.co/datasets/biglam/yalta_ai_tabular_dataset.hw1-tabular-hand-data
24-679 (Fall 2026): CMU Right-Hand Measurements
cmuchancel/hw1-tabular-hand-data
Right-hand finger measurements from Carnegie Mellon University students, plus explicitly marked
synthetic training variants. The classroom regression task predicts middle-finger length from thumb,
index-finger, ring-finger, and pinky lengths together with the recorded Female / Male categorical
feature. All finger measurements are stored in centimeters.
Source and task
The original… See the full description on the dataset page: https://huggingface.co/datasets/cmuchancel/hw1-tabular-hand-data.2026-24679-tabular-dataset
24-679 (Fall 2026): Music Listening Survey
ccm/2026-24679-tabular-dataset
Course-survey responses about music listening, plus explicitly marked synthetic training variants.
The classroom regression task predicts weekly listening hours from two count features and five
categorical preferences. Stored haiku answers provide source context and are excluded from this
task's predictors.
Source and task
The preparation notebook reads 24-679-tabular-survey.csv, removes the… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2026-24679-tabular-dataset.2026-24679-tabular-dataset
24-679 (Fall 2026): Snack Nutrition Dataset
kwongnon/2026-24679-tabular-dataset
Nutrition-label information for 50 packaged snack products. Each row represents one snack and
contains its product name, servings per container, calories, total fat, cholesterol, sodium,
total carbohydrate, and protein.
The dataset can be used for classroom exercises involving tabular data exploration, preprocessing,
visualization, clustering, regression, or other machine-learning tasks using… See the full description on the dataset page: https://huggingface.co/datasets/kwongnon/2026-24679-tabular-dataset.foods-tabular-datasetcar-classification-tabular-datasettest_composite_and_tabular_datasetdatabase-small-tabular-regressionhw1-24679-tabular-dataset
Shoe Size Measurements Tabular Dataset
Dataset Summary
Purpose: This dataset was created for tabular data analysis and prediction tasks involving shoe measurements, developed as part of CMU 24-679 coursework to explore tabular data augmentation techniques.
Quick Stats:
338 total samples (30 original + 308 augmented)
3 numerical features + 3 categorical features
High correlation between size measurements (>0.97)
~10x augmentation factor
Contact: maryzhang@cmu.edu… See the full description on the dataset page: https://huggingface.co/datasets/maryzhang/hw1-24679-tabular-dataset.database-small-tabular-regression22026-24679-tabular-dataset
24-679 (Fall 2026): Everyday Object Measurements
kadireks/2026-24679-tabular-dataset
Hand-measured everyday objects from one household, plus explicitly marked synthetic training
variants. The classroom classification task predicts an object's category from three ruler
measurements and two categorical properties. The object's own name is stored as source context and is
excluded from this task's predictors.
Source and task
The preparation notebook reads objects.csv… See the full description on the dataset page: https://huggingface.co/datasets/kadireks/2026-24679-tabular-dataset.database-small-tabular-regression4Jobs-tabular-datasetJobs-tabular-datasetbooks-tabular-dataset
📄 Model Card: Books Tabular Dataset
1. Purpose
This dataset was created for educational purposes in the context of Homework 1 (Dealing with Data). The goal is to provide a small but structured tabular dataset that allows students to practice working with real-world features, preprocessing, augmentation, and uploading to Hugging Face. The dataset supports tasks such as classification, exploratory data analysis (EDA), and simple modeling.
2. Composition… See the full description on the dataset page: https://huggingface.co/datasets/EricCRX/books-tabular-dataset.train_composite_and_tabular_datasetBooks-tabular-dataset2025-24679-pens-tabular-dataset
Pens Dataset
This is a simple, manually-collected tabular dataset containing the physical attributes of 33 unique writing pens. It is designed for straightforward binary classification tasks. The features include quantitative measurements like length, and qualitative features like ink color, body material, and brand.
The primary target variable is cap_presence, a binary indicator (1 or 0) signifying whether a pen has a separate cap or not (e.g., a capped pen vs. a retractable… See the full description on the dataset page: https://huggingface.co/datasets/zacCMU/2025-24679-pens-tabular-dataset.2025-24679-tabular-datasetHW1-tabular-dataset
Dataset Card for Book Tabular Data
This tabular dataset provides measurements on books selected from my bookshelf.
Dataset Details
Dataset Description
For a selection of books on my bookshelf, I collected some tabular data.
I selected 15 fiction and 15 nonfiction books.
I then documented how many pages each had, how thick the book was, if I had read it/ started it/ not read it, and if it was a book I would recommend to everyone.
These variables were… See the full description on the dataset page: https://huggingface.co/datasets/jennifee/HW1-tabular-dataset.2025-24679-tabular-dataset
Dataset Card for ccm/2025-24679-tabular-dataset
This dataset captures self-reported music listening behaviors and preferences, including listening hours, library size, playlist creation, sharing habits, preferred decades, and social context (alone vs. with others). It was created as a class exercise in data collection and augmentation.
Dataset Details
Dataset Description
This dataset captures self-reported music listening behaviors and preferences, including… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2025-24679-tabular-dataset.HW1-augmented-tabular-datasetL1_tabular_data
Dataset Card for "L1_tabular_data"
More Information needed
2025-24679-tabular-Kita-decode-dataset
Dataset Card for Doors
This is a dataset of doors in our everday life.
Uses
No specific use.
Dataset Structure
Just tabular
Dataset Creation
Source Data
Temu, Schools
hw1-24679-tabular-dataset-resulttabular-data-testpens-markers-tabular-dataset-2025
Dataset Card for keerthikoganti/pens-markers-tabular-dataset-2025
Dataset Details
Dataset Description
This dataset contains measurements and categorical attributes of common pens like Pilot, Bic, Sharpie. It was created as a class exercise for supervised learning on tabular data, supporting both regression by predicting line width and classification such as “thick vs. thin” line.
Curated by: Fall 2025 24-679 course at Carnegie Mellon University
Shared by :… See the full description on the dataset page: https://huggingface.co/datasets/keerthikoganti/pens-markers-tabular-dataset-2025.tabular_dataset
