tabular-data
po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabular-logs-and-datasetstabular-data-wf9uh
Tabular Data Wf9Uh
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
3,251
Validation
409
Test
206
Total
3,866
Classes (12)
bold_parent_row
bold_row
closure_row
column
direct_children
non_bold_parent_row
non_bold_row
parent_column
prime_parent
sub_row
table
Usage
With LibreYOLO
from libreyolo… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/tabular-data-wf9uh.multitask-tabular-datasetsThis is a port of the Multi-Label Classification Dataset Repository (link).
We convert the datasets from there to simple csvs, resulting in 32 csvs (many of their mulan files fail to parse into python for us)
The targets in each csv are labeled with the suffix __target
Dataset
Domain
m
d
q
Card
Dens
Div
avgIR
rDep
m×q×d
0
3s-bbc1000
Text
352
1000
6
1.125
0.188
0.234
1.718
0.733
2.11e+06
1
3s-guardian1000
Text
302
1000
6
1.126
0.188
0.219
1.773
0.667
1.81e+06
2
3s-inter3000… See the full description on the dataset page: https://huggingface.co/datasets/imodels/multitask-tabular-datasets.IFT-Data-For-Tabular-Tasksyash-gym-tabular-dataset
Yash Gym Tabular Dataset
Dataset Summary
This dataset contains information on 30 unique gym machines with 5 consistent features and a binary target (Upper/Lower).It includes:
original: 30 manually collected samples
augmented: ~300 synthetic samples created with jitter, SMOTE-NC, MixUp, and CTGAN.
Intended Use
Educational dataset for tabular ML tasks, demonstrating preprocessing + augmentation.Not suitable for prescribing exercise or medical advice.… See the full description on the dataset page: https://huggingface.co/datasets/ysakhale/yash-gym-tabular-dataset.
