datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IntTravel_datasetAll data has been uploaded. Please note that the current POI dataset does not include coordinate information. We will be updating it as soon as possible, so please stay tuned. If you do not need the coordinate information for the POIs, please disregard this message.
IntTravel: A Real-World Dataset and Generative Framework for Integrated Multi-Task Travel Recommendation
Paper | GitHub
IntTravel is the first large-scale public dataset for integrated travel recommendation, including… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/IntTravel_dataset.trinity-dataset-v3rtl-ml-dataset
RTL-ML Dataset v2
Dataset Summary
This dataset contains 800 validated RF signal samples captured using an RTL-SDR Blog V4 dongle on an Indiedroid Nova (RK3588S). Designed for training machine learning models to classify common RF signals.
Samples: 800 (7 classes)
Format: NumPy arrays (.npy files) — each file is a dict with IQ data + metadata
Sample Rate: 1.024 MSPS
Sample Duration: 0.5 seconds per capture
Quality Gates: DC removal, auto-gain, 6 dB minimum SNR, per-class… See the full description on the dataset page: https://huggingface.co/datasets/TrevTron/rtl-ml-dataset.rtl-ml-dataset
RTL-ML Dataset
Dataset Summary
This dataset contains 240 validated RF signal samples captured using an RTL-SDR Blog V4 dongle. It's designed for training machine learning models to classify common RF signals.
Total Size: 1.9 GBSamples: 240 (30 samples × 8 classes)Format: NumPy arrays (.npy files)Sample Rate: 1.024 MSPSSample Duration: 1 second per capture
Signal Classes
Class
Frequency
Count
Description
ADS_B
1090 MHz
30
Aircraft transponder… See the full description on the dataset page: https://huggingface.co/datasets/weswebunited/rtl-ml-dataset.dataset_concrete_compressive_strength
Machine learning in concrete science: applications, challenges, and best practices
Dataset containing concrete compressive strength for 1030 materials
Dataset Information
Source: Foundry-ML
DOI: 10.18126/8k1f-mx77
Year: 2022
Authors: Li, Zhanzhao, Yoon, Jinyoung, Zhang, Rui, Rajabipour, Farshad, Srubar III, Wil V., Dabo, Ismaila, Radlińska, Aleksandra
Data Type: tabular
Fields
Field
Role
Description
Units
Cement (component 1)(kg in a m^3… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ml/dataset_concrete_compressive_strength.vFab-1.0-Photolithography-ML-surrogate-Tensor-Dataset
vFab-1.0-Photolithography-ML-surrogate-Tensor-Dataset
This dataset contains generated physics simulations for semiconductor photolithography, formatted as Native NCHW .npy arrays for direct ingestion into PyTorch models.
Dataset Structure
The data is split into two simulation phases:
1. Aerial Image Dataset (aerial_image_dataset/)
Inputs (X_input/): Shape [4, 512, 512]. Contains:
Channel 0: Binary Layout Mask (Vertical/Horizontal Gratings, Contact Arrays, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/nk-21-mit/vFab-1.0-Photolithography-ML-surrogate-Tensor-Dataset.ML-Proto-DatasetSynthetic-AI-ML-Dataset
Synthetic-AI-ML-Dataset
Synthetic Q&A dataset on AI and Machine Learning
Dataset Details
Metric
Value
Topic
AI and Machine Learning
Total Q&A Pairs
14021
Valid Pairs
14021
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
14,941,957
Completion Tokens
17,159,263
Total Tokens
32,101,220
GPU Energy
12.9628 kWh
Sources
This dataset was generated from 474 scholarly papers:
#… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.single_objectrobosuite_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 875,
"total_frames": 12575,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:875"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/ML-GOD/single_objectrobosuite_dataset.kknews-dataset
KKNews.uz Dataset
Qaraqalpaqstan Xabar Agentligi (kknews.uz) maqalaları — 5 tilde.
Languages
Code
Language
ru
Russian
uz
Uzbek (Latin)
oz
Uzbek (Cyrillic)
kk
Karakalpak (Cyrillic)
qq
Karakalpak (Latin)
Columns
Column
Type
Description
id
int
WordPress post ID
lang
string
Language code
category_id
int
Category ID
category_name
string
Category name
title
string
Plain text title
content_html
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/kknews-dataset.ML-Music-Classifier-dataset-and-model-name-Models
🎧 Spotify Music Preference Analysis
🧠 Project Overview
This project analyzes Spotify music data to predict song preferences using machine learning models. The analysis is based on a dataset of 195 songs (100 liked, 95 disliked) with various audio features extracted from Spotify's API.
📂 Dataset Description
📥 Data Collection Process
Liked Songs (100 tracks):
🎵 Primarily French Rap
🎸 Some American Rap, Rock, and Electronic music
✅… See the full description on the dataset page: https://huggingface.co/datasets/Jack1808/ML-Music-Classifier-dataset-and-model-name-Models.deprem_satellite_semantic_whu_dataset
Dataset Card for "deprem_satellite_semantic_whu_dataset"
More Information needed
PrahaTTS-ML-Expressive-Datasetshipaker-dataset
Shipaker.uz Dataset
A collection of health and medicine articles in the Karakalpak language, scraped from shipaker.uz. The site is run by a public health organization in Karakalpakstan, Uzbekistan, and publishes articles on topics such as disease prevention, nutrition, psychology, and general wellness.
Columns
Column
Type
Description
id
int
Article ID
title
string
Article title
content_html
string
Full article body (original HTML)
content_text… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/shipaker-dataset.support-routing-ml-20260910-dataset
Applied ML Support Router Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Support operations need reproducible routing models that expose confidence and defer uncertain cases.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/support-routing-ml-20260910-dataset.ai-ml-instruction-dataset
AI/ML Engineering Instruction Dataset
Comprehensive instruction dataset covering machine learning concepts, PyTorch implementations, NLP with transformers, model evaluation, and feature engineering.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Ai Ml topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/ai-ml-instruction-dataset.ml-tutor-datasetML_EdgeIIoT_datasetextended_TBE_datasetdataset_metallicglass_dmax
Metallic Glasses and their Properties
Dataset containing experimental max casting diameters of 998 metallic glasses
Dataset Information
Source: Foundry-ML
DOI: 10.18126/fs5e-kr15
Year: 2021
Authors: Voyles, Paul M, Schultz, Lane E., Morgan, Dane, Francis, Carter, Afflerbach, Benjamin, Hakeem, Abdulrhman
Data Type: tabular
Fields
Field
Role
Description
Units
Composition
input
Material composition
Reference
input
Original data reference… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ml/dataset_metallicglass_dmax.COIN-PHP-YOLOv11n-Dataset
Coin Optical Identification and Numeration for the Philippine Peso - Dataset
About
This repo contains the dataset i gathered, annotated, and processed to train a yolov11n model for detecting and counting philippine peso coins.
Folder Structure
I organized the repo into three main folders to keep things clean:
/raw - contains the original unedited photos.
/annotated - contains the images and their respective label files from the annotation tool… See the full description on the dataset page: https://huggingface.co/datasets/ml-remunn/COIN-PHP-YOLOv11n-Dataset.clickbait-ml_datasetml-interview-sft-dataset
ML/AI Interview Coach — SFT Dataset
A curated dataset of 566 high-quality Q&A pairs covering ML, Deep Learning, NLP, LLMs, RAG, Vector Databases, LangChain, Agentic AI, MLOps, and more — designed for fine-tuning an ML Interview Coach model.
Dataset Summary
Stat
Value
Total Q&A pairs
566
Unique topics
75
Format
ChatML (system + user + assistant)
Language
English
Avg answer length
~800 tokens
Sources
15+ interview prep documents + hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/raghu298/ml-interview-sft-dataset.support-routing-ml-20260920-dataset
Applied ML Support Router Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Support operations need reproducible routing models that expose confidence and defer uncertain cases.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/support-routing-ml-20260920-dataset.ml-algorithm-dataset
ml-algorithm-dataset
A conjecture of datasets specifically designed for Machine Learning training and tuning pipelines, mostly novel algorithms and their representations as RAW ASCII and LaTeX, connected to the asi-ecosystem framework.
dataset_mg_alloy
Prediction of mechanical properties of biomedical magnesium alloys based on ensemble machine learning
Dataset containing mechanical properties of 365 Mg alloys
Dataset Information
Source: Foundry-ML
DOI: 10.18126/myj4-0h48
Year: 2023
Authors: Hou, Haobing, Wang, Jianfeng, Ye, Li, Zhu, Shijie, Wang, Liguo, Guan, Shaokang
Data Type: tabular
Fields
Field
Role
Description
Units
Mg(wt.%)
input
Amount of Mg
wt%
Zn(wt.%)
input
Amount of Zn
wt%
Y(wt.%)… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ml/dataset_mg_alloy.dataset_rpv_tts
Predictions and uncertainty estimates of reactor pressure vessel steel embrittlement using Machine learning
Dataset containing 4535 transition temperature shifts of reactor pressure vessel steels
Dataset Information
Source: Foundry-ML
DOI: 10.18126/3zkm-yd51
Year: 2023
Authors: Jacobs, Ryan, Yamamoto, Takuya, Odette, G. Robert, Morgan, Dane
Data Type: tabular
Fields
Field
Role
Description
Units
temperature_C
input
Temperature of measurement
degC… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ml/dataset_rpv_tts.oxide-electrocatalyst-ml-datasetfile:///app/README.md
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset("ANBU963/oxide-electrocatalyst-ml-dataset")
AI4C-ML-Dataset
Data Card for Dataset of "Autonomous discovery of atomically dispersed catalysts from a ten-million-scale configuration space"
This is a dataset repository. Please check all the information in our paper and GitHub page.
Get Start
GPGB_AL
H2O2_decom_bar
Check this code for using the dataset.
Licensing Information
This dataset is licensed under the Apache license 2.0.
header_dataset_24_06
