datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabular-benchmark
Tabular Benchmark
Dataset Description
This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms.
Repository: https://github.com/LeoGrin/tabular-benchmark/community
Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document
Dataset Summary
Benchmark made of curation of various tabular data learning tasks, including:
Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.demo-tabular-benchmarks
📊 Carla HQ Tabular Foundation Model Benchmarks
Centralized benchmark repository of canonical tabular datasets curated for Carla HQ and TabICL (In-Context Learning foundation models for tabular data).
Each dataset is hosted as an independent subset/config with native Parquet storage, schema qualities, OpenML source links, and synchronized Google Sheets for live spreadsheet experimentation.
🚀 Quickstart & Download Options
Option 1: Using… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmarks.CausalArena
CausalArena public release
This repository contains the public CausalArena dataset release: executable SCMs, selected result tables, and real-data source indices.
What is included
scm/: the public half of each generated SCM family: 500 synthetic SCM configurations, 50 semantic SCMs, and 50 formula-grounded SCMs. Released SCMs include both observation-only and observation-plus-intervention exports.
scm/{semantic,formula}/artifacts/: per-scenario graph, generator… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/CausalArena.tcga-tgct-tabular-open
TCGA-TGCT — Tabular (Open Access)
Open-access TCGA-TGCT data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:21:25 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-tgct-tabular-open.Artistic_Landscape
Artistic Landscape Dataset
Welcome to the Artistic Landscape.
Overview
Artistic Landscape is a curated collection of synthetically generated imagery designed to explore conceptual combinations across a wide variety of art styles, media, and visual characteristics. It is intended to be exploratory in scope, comparative across many aesthetic dimensions, and practical for downstream research workflows where a large, structured visual vocabulary is useful.
In total, we used… See the full description on the dataset page: https://huggingface.co/datasets/tabularisai/Artistic_Landscape.tcga-acc-tabular-open
TCGA-ACC — Tabular (Open Access)
Open-access TCGA-ACC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:44:59 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-acc-tabular-open.tcga-read-tabular-open
TCGA-READ — Tabular (Open Access)
Open-access TCGA-READ data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:16:05 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-read-tabular-open.demo-tabular-benchmark-containers
📦 Carla HQ — Tabular Benchmark CuratedContainers
Centralized repository of Data Foundry CuratedContainers curated for Carla HQ, TabICLv2, and the next generation of Tabular Foundation Models (TabPFN, EXAONE, Google TabFM).
Each container directory provides:
Columnar Parquet Data (dataset.parquet): Clean, type-normalized, and validated tabular dataset binary.
Standardized Task Molds (task_metadata.predictive-ml-task-mold-v1.json): Problem definitions, target attributes, and… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmark-containers.tcga-kich-tabular-open
TCGA-KICH — Tabular (Open Access)
Open-access TCGA-KICH data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:58:23 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-kich-tabular-open.tcga-dlbc-tabular-open
TCGA-DLBC — Tabular (Open Access)
Open-access TCGA-DLBC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:53:50 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-dlbc-tabular-open.tcga-paad-tabular-open
TCGA-PAAD — Tabular (Open Access)
Open-access TCGA-PAAD data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:12:24 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-paad-tabular-open.tcga-uvm-tabular-open
TCGA-UVM — Tabular (Open Access)
Open-access TCGA-UVM data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:27:20 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-uvm-tabular-open.tcga-laml-tabular-open
TCGA-LAML — Tabular (Open Access)
Open-access TCGA-LAML data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:02:09 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-laml-tabular-open.tcga-pcpg-tabular-open
TCGA-PCPG — Tabular (Open Access)
Open-access TCGA-PCPG data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:13:13 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-pcpg-tabular-open.tcga-ucs-tabular-open
TCGA-UCS — Tabular (Open Access)
Open-access TCGA-UCS data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:27:02 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-ucs-tabular-open.tcga-chol-tabular-open
TCGA-CHOL — Tabular (Open Access)
Open-access TCGA-CHOL data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:38:00 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-chol-tabular-open.tcga-thym-tabular-open
TCGA-THYM — Tabular (Open Access)
Open-access TCGA-THYM data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:24:24 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-thym-tabular-open.tcga-esca-tabular-open
TCGA-ESCA — Tabular (Open Access)
Open-access TCGA-ESCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:54:06 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-esca-tabular-open.oak
NEWS:
A new version of the dataset with 120,000,000 more tokens is upload: OAK v1.1
Open Artificial Knowledge (OAK) Dataset
Overview
The Open Artificial Knowledge (OAK) dataset is a large-scale resource of over 650 Millions tokens designed to address the challenges of acquiring high-quality, diverse, and ethically sourced training data for Large Language Models (LLMs). OAK leverages an ensemble of state-of-the-art LLMs to generate high-quality text… See the full description on the dataset page: https://huggingface.co/datasets/tabularisai/oak.tcga-meso-tabular-open
TCGA-MESO — Tabular (Open Access)
Open-access TCGA-MESO data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:10:59 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-meso-tabular-open.bike-sharing-tabular
Bike Sharing Demand - Hourly (Poisson)
A ready-to-use copy of the UCI Bike Sharing Dataset (hourly granularity,
17,379 × 17), accompanied by baseline metrics from an 8-architecture tabular
modelling pipeline for direct comparison.
Originally collected and published by Fanaee-T & Gama (2014). Source:
UCI ML Repository id 275.
At a glance
Field
Value
Rows
17,379 hourly observations
Time range
Jan 2011 - Dec 2012
Columns
17 (16 features + 1 target)… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/bike-sharing-tabular.tcga-stad-tabular-open
TCGA-STAD — Tabular (Open Access)
Open-access TCGA-STAD data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:19:42 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-stad-tabular-open.tcga-kirc-tabular-open
TCGA-KIRC — Tabular (Open Access)
Open-access TCGA-KIRC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:58:45 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-kirc-tabular-open.tcga-gbm-tabular-open
TCGA-GBM — Tabular (Open Access)
Open-access TCGA-GBM data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:54:54 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-gbm-tabular-open.yash-gym-tabular-dataset
Yash Gym Tabular Dataset
Dataset Summary
This dataset contains information on 30 unique gym machines with 5 consistent features and a binary target (Upper/Lower).It includes:
original: 30 manually collected samples
augmented: ~300 synthetic samples created with jitter, SMOTE-NC, MixUp, and CTGAN.
Intended Use
Educational dataset for tabular ML tasks, demonstrating preprocessing + augmentation.Not suitable for prescribing exercise or medical advice.… See the full description on the dataset page: https://huggingface.co/datasets/ysakhale/yash-gym-tabular-dataset.tcga-coad-tabular-open
TCGA-COAD — Tabular (Open Access)
Open-access TCGA-COAD data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:52:16 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-coad-tabular-open.tcga-blca-tabular-open
TCGA-BLCA — Tabular (Open Access)
Open-access TCGA-BLCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:45:21 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-blca-tabular-open.tabular-errors-v1
TabFix multilingual table error pairs — version 2.0
This release keeps 18 business error categories and separates executable deterministic detection from two residual neural categories: text.encoding and text.spelling. The same repository and family-disjoint splits are retained.
Split
Records
Open-vocabulary views
train
27948
3260
validation
17127
1844
test
32776
3540
The seven string columns remain id, split, family_id, clean_xml, corrupt_xml, errors… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/tabular-errors-v1.
