datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
suno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.CostNav-Teleop-Dataset
CostNav Teleop Dataset
Dataset Summary
The CostNav Teleop Dataset is a large-scale collection of human teleoperation recordings for robot navigation in an urban sidewalk simulation environment. It was collected as part of the CostNav benchmark, which evaluates navigation systems using real-world economic cost and revenue metrics rather than purely technical metrics.
The dataset contains 2,203 teleoperation episodes totaling 50.2 hours of driving… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/CostNav-Teleop-Dataset.AniGen-Sample-Dataset
AniGen Sample Data
This directory is a compact example subset of the AniGen training dataset.
What Is Included
10 examples
10 unique raw assets
Full cross-modal files for each example
A subset metadata.csv with 10 rows
The retained directory layout follows the core structure of the reference test set:
raw/
renders/
renders_cond/
skeleton/
voxels/
features/
metadata.csv
statistics.txt
latents/ (encoded by the trained slat auto-encoder)
ss_latents/ (encoded by the… See the full description on the dataset page: https://huggingface.co/datasets/VAST-AI/AniGen-Sample-Dataset.DuET-dataset
DuET TE measurements of 64 human cell types & benchmark datasets for TE and MRL prediction task
This dataset comprises TE datasets for 64 cell types and benchmark datasets for TE and MRL prediction task.
How to setup
First, clone the main repository to your work directory:
$ git clone https://github.com/mogam-ai/DuET.git
$ cd DuET
Then, download the dataset repository into DuET/datasets subdirectory.
# Needs huggingface-cli (pip install huggingface-cli)
$ huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/DuET-dataset.chinese-ai-detection-dataset
Chinese AI Detection Dataset
中文AI文本检测数据集
数据集简介
用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。
核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。
数据统计
类型
样本数
说明
总计
66,001
训练/验证/测试集
纯人类
27,719
多领域人类文本
纯AI
27,719
多模型生成
C2 (续写)
3,781
人类开头+AI续写
C3 (改写)
3,781
AI改写人类文本
C4 (润色)
3,001
AI润色人类文本
数据格式
{
"text": "文本内容(混合文本包含[SEP]标记)",
"label": 0, // 0=Human, 1=AI
"category": "C2", // Human/AI/C2/C3/C4
"source": "数据来源"
}… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.hotel_datasetsAI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.AI_financial_fraud_datasetkrakow-handover-dataset
FUMD-AI Kraków Vehicular Handover Dataset
AI-ready, per-second traces of vehicles moving through a simulated 5G NR
deployment over central Kraków, with every serving-cell change classified and
every row labelled with an upcoming-handover flag and the cell the vehicle will
settle on.
Each row joins a vehicle's mobility state (position, speed, heading, lane)
to its radio state at the same second (serving cell, SINR, CQI, RLC
delay/throughput, distance to the serving gNB). The… See the full description on the dataset page: https://huggingface.co/datasets/fumd-ai/krakow-handover-dataset.microduck-locomotion-dataset-v1
🦆 MicroDuck Locomotion Dataset V1.0: Zero-Slip Flat Ground Baseline
Publisher: devorah-ai-2026 / devorahai2026License: Commercial / CC-BY-SA 4.0Task: Bipedal Locomotion, Reinforcement Learning, Sim-to-RealHardware Platform: Antigravity 14-DOF MicroDuck Biped Robot
📌 Overview (개요)
본 데이터셋은 **14-DOF 소형 2족 보행 로봇(MicroDuck)**의 완벽한 평면 보행(Flat-ground walking) 궤적 및 제어 로그를 담은 고품질 합성 데이터(Synthetic Data)입니다. 기존 강화학습 기반 2족 보행 에이전트들이 흔히 겪는 발바닥 미끄러짐(Ice-Skating) 및… See the full description on the dataset page: https://huggingface.co/datasets/devorah-ai-2016/microduck-locomotion-dataset-v1.enterprise-financial-crime-ai-datasetTransactions → Risk Analysis → Alerts → Investigation → SAR Reports
Dataset Statistics
Total records: 310,396Dataset size: 339 MBAuto-converted parquet size: 65 MB
Languages:
English
French
Spanish
Main fields:
email_id
thread_id
timestamp
language
bank
department
country
risk_level
Enterprise Financial Crime AI Dataset
The Enterprise Financial Crime AI Dataset is a high-fidelity dataset built from real-world operational patterns and enterprise data structures… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/enterprise-financial-crime-ai-dataset.time_series_datasets
Tourism Monthly Time Series Dataset with Economic and Static Covariates
This dataset, originally sourced from Athanasopoulos et al. (2011), focuses on the tourism industry with a monthly frequency and has been enhanced with economic covariates (e.g., CPI, Inflation Rate, GDP) from official Australian government sources. We also perform some preprocessing to further increase the usability of the dataset with dynamic start dates for each series and static covariates for in-depth time… See the full description on the dataset page: https://huggingface.co/datasets/zaai-ai/time_series_datasets.alicante-handover-dataset
FUMD-AI Alicante Vehicular Handover Dataset
Per-second traces of vehicles moving through a simulated 5G NR deployment over
central Alicante, Spain, with every serving-cell change classified and every row
labelled with an upcoming-handover flag and the cell the vehicle will settle on.
Each row joins a vehicle's mobility state (position, speed, heading, lane) to
its radio state at the same second (serving cell, SINR, CQI, RLC
delay/throughput, distance to the serving gNB).
This is… See the full description on the dataset page: https://huggingface.co/datasets/fumd-ai/alicante-handover-dataset.AI_financial_fraud_datasetProjectManagementLLM_datasetProject Management Data (Synthetic Dataset)
This repository contains a synthetic dataset representing project performance data. The dataset includes details of various tasks within different projects, such as cost variance, schedule variance, and task descriptions. It is meant to simulate the type of data typically used in project management performance analysis.
Dataset Description
The dataset is structured to contain information about multiple projects, each with a set of tasks. For each… See the full description on the dataset page: https://huggingface.co/datasets/ai-in-projectmanagement/ProjectManagementLLM_dataset.Ai_ethics_dataset
AI Ethics Preference Annotation Dataset
A human-annotated preference dataset for RLHF and Direct Preference Optimization (DPO), focused on AI ethics failure modes. 95 prompts, 190 response pairs, full annotation across five dimensions.
Annotator: Mandy Hathaway — AI ethics specialist and technical writer with an MA in Ethical Technology & Artificial Intelligence. mandyhathaway.com
Dataset Summary
Most public preference datasets optimize for general helpfulness or… See the full description on the dataset page: https://huggingface.co/datasets/animasuri/Ai_ethics_dataset.LLM01-datasetsample-structure-dataset
Sample dataset for PETAL model
This dataset is a sample dataset to test the functionalities of the PETAL model (encoder and decoder).
It is based on CASP15 dataset, see
https://predictioncenter.org/casp15/
https://github.com/Bhattacharya-Lab/CASP15
The registries folder contains the registry of CASP15 dataset (a csv file with filename, pdb_id, etc.)
abidhussai512_ai-tools-usage-and-trends-dataset-2026
🚀 AI Tools Usage & Trends Dataset (2026)
AI Tools Usage & Trends Dataset Mapping the Explosive Growth of Artificia
Dataset Info
Source: Kaggle
Original Size: 0.33 MB
Kaggle Downloads: 47
Files: 1
Files
AI_tools_dataset.csv
Mirrored from Kaggle
ELSA-Emotion-and-Language-Style-Alignment-Dataset
ELSA: Emotion and Language Style Alignment Dataset
The ELSA (Emotion and Language Style Alignment) dataset provides fine-grained emotional rewrites of text across four stylistic contexts: conversational, formal, poetic, and narrative. It is designed to support research in emotion-conditioned generation, stylistic variation, and affect-aware NLP.
Overview
Source: Based on the dair-ai/emotion dataset and emotion labels aligned with the GoEmotions taxonomy.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/joyspace-ai/ELSA-Emotion-and-Language-Style-Alignment-Dataset.uk_retail_store_synthetic_dataset
Synthetic Data Generation Demo — UK Retail Dataset
Welcome to this synthetic data generation demo repository by Syncora.ai. This project showcases how to generate synthetic data using real-world tabular structures, demonstrated on a UK retail dataset with columns such as:
Country
CustomerID
UnitPrice
InvoiceDate
Quantity
StockCode
This dataset is designed for dataset for LLM training and AI development, enabling developers to work with privacy-safe, high-quality… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/uk_retail_store_synthetic_dataset.financial_credit_dataset
Financial Credit Card Dataset — Free Financial Dataset 💳
High-Fidelity Financial Dataset for ML & AI Research, Credit Risk Modeling, and LLM Training
🌟 About This Dataset
This financial dataset provides synthetic credit and debit card records, including card brand, type, credit limits, issuance dates, CVV, and more.All records are privacy-safe, making it ideal for ML experimentation, AI research, and dataset for LLM training.
Visit our website to learn more about the… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/financial_credit_dataset.DataSet-Of-using-AI
Dataset of Using AI
Team Members
Yusif Aziz Ibrahim
Yousif Kawa Faris
Ahmed Qaraman
Overview
This dataset shows adoption rates of popular AI tools across age groups.It includes usage values for ChatGPT, Gemini, DeepSeek, Claude, and Nova.Data was manually compiled and organized in Excel.
Structure
Age Groups: 13–17, 18–24, 25–34, 35–44, 45–54, 55+
Models: ChatGPT, Gemini, DeepSeek, Claude, Nova
Values: Usage rates… See the full description on the dataset page: https://huggingface.co/datasets/yusifhajiaziz/DataSet-Of-using-AI.AI-and-Data-Science-Job-Market-Dataset
About Dataset
Edit
The AI & Data Science Job Market Dataset (2020–2026) is a synthetically generated dataset designed to simulate real-world hiring patterns across the artificial intelligence and data science job market.
The dataset contains structured information about job roles, company characteristics, required technical skills, education levels, experience requirements, and salary ranges. It reflects hiring data across multiple countries, industries, and company sizes.
This… See the full description on the dataset page: https://huggingface.co/datasets/shree0910/AI-and-Data-Science-Job-Market-Dataset.time_series_dataset_residuals
Tourism Monthly Time Series Dataset with Economic and Static Covariates
This dataset, originally sourced from Athanasopoulos et al. (2011), focuses on the tourism industry with a monthly frequency and has been enhanced with economic covariates (e.g., CPI, Inflation Rate, GDP) from official Australian government sources. We also perform some preprocessing to further increase the usability of the dataset with dynamic start dates for each series and static covariates for in-depth time… See the full description on the dataset page: https://huggingface.co/datasets/zaai-ai/time_series_dataset_residuals.Ai_ethics_dataset
AI Ethics Preference Annotation Dataset
license: cc-by-4.0
task_categories:
text-generation
text-classification
task_ids:
language-modeling
tags:
rlhf
dpo
preference-learning
ai-ethics
ai-safety
alignment
human-feedback
annotation
language:
en
size_categories:
n<1K
pretty_name: AI Ethics Preference Annotation Dataset
A human-annotated preference dataset for RLHF and Direct Preference Optimization (DPO), focused on AI ethics failure modes. 95 prompts, 190 response pairs, full… See the full description on the dataset page: https://huggingface.co/datasets/philosophyFire/Ai_ethics_dataset.Scouter-ai-dataset
Scouter AI Dataset
This dataset contains soccer scouting reports and player metrics.
Dataset Schema
Column
Type
Name
string
Age
int64
Position
string
Team
string
League
string
Nation
string
Rating
int64
Market_Value
string
Scout_Report
string
⚽ AI-Powered Synthetic Soccer Scout Dataset
Overview
This project generates a high-quality synthetic dataset of 20,000 professional soccer players using Python, Faker, and a… See the full description on the dataset page: https://huggingface.co/datasets/yonaitay/Scouter-ai-dataset.algozee_ai-driven-global-market-intelligence-dataset
AI-Driven Global Market Intelligence Dataset
Global Financial Market Data for Risk, Trend, and Investment Analysis
Dataset Info
Source: Kaggle
Original Size: 9.04 MB
Kaggle Downloads: 166
Files: 1
Files
global_market_ai_dataset.csv
Mirrored from Kaggle
Ai_ethics_dataset
AI Ethics Preference Annotation Dataset
license: cc-by-4.0
task_categories:
text-generation
text-classification
task_ids:
language-modeling
tags:
rlhf
dpo
preference-learning
ai-ethics
ai-safety
alignment
human-feedback
annotation
language:
en
size_categories:
n<1K
pretty_name: AI Ethics Preference Annotation Dataset
A human-annotated preference dataset for RLHF and Direct Preference Optimization (DPO), focused on AI ethics failure modes. 95 prompts, 190 response pairs, full… See the full description on the dataset page: https://huggingface.co/datasets/Emilynnjk/Ai_ethics_dataset.
