datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
auto-mpg
Auto Miles per Gallon (MPG) Dataset
Following description was taken from UCI machine learning repository.
Source: This dataset was taken from the StatLib library which is maintained at Carnegie Mellon University. The dataset was used in the 1983 American Statistical Association Exposition.
Data Set Information:
This dataset is a slightly modified version of the dataset provided in the StatLib library. In line with the use by Ross Quinlan (1993) in predicting the attribute… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/auto-mpg.iac-eval
IaC-Eval dataset (v1.1)
IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities.
This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now).
| Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper |
2. Usage instructions
Option 1: Running the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/autoiac-project/iac-eval.Auto-ACD
Auto-ACD
Auto-ACD is a large-scale, high-quality, audio-language dataset, building on the prior of robust audio-visual correspondence in existing video datasets, VGGSound and AudioSet.
Homepage: https://auto-acd.github.io/
Paper: https://huggingface.co/papers/2309.11500
Github: https://github.com/LoieSun/Auto-ACD
Analysis
Auto-ACD, comprising over 1.9M audio-text pairs.
As shown in figure, The text descriptions in Auto-ACD contain long texts (18 words) and… See the full description on the dataset page: https://huggingface.co/datasets/Loie/Auto-ACD.autonlp-data-peptidesDeep learning the collisional cross sections of the peptide universe from a million experimental values
Data generated from MaxQuant output
wget https://ftp.pride.ebi.ac.uk/pride/data/archive/2020/12/PXD017703/HeLa_200ng_Library_MaxQuant.zip
unzip HeLa_200ng_Library_MaxQuant.zip
awk -F '\t' '{print $1,",",$40}' evidence.txt > pepCCS.csv
wc pepCCS.csv
352111 1056333 12736697 pepCCS.csv
Code
nier-automata-wiki-dataset
NieR: Automata Wiki Dataset
A comprehensive, structured knowledge base extracted from NieR: Automata wikis, optimized for machine learning, RAG systems, and AI applications.
📊 Dataset Overview
This dataset contains cleaned, structured knowledge from the NieR: Automata Fandom Wiki, covering:
Characters: YoRHa units (2B, 9S, A2), Machine Lifeforms, NPCs
Locations: City Ruins, Desert Zone, Bunker, Forest Kingdom
Technology: Pod Programs, Weapons, Black Box systems… See the full description on the dataset page: https://huggingface.co/datasets/Lvoxx/nier-automata-wiki-dataset.autoencoder-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models.
It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses).
The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.autonlp-data-mami-semeval-20-21autonlp-data-CoronaIt's all about Corona
data-preprocessing-automl-benchmarks
Data Preprocessing AutoML Benchmarks
This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML.
Usage
Load a specific dataset configuration like this:
from datasets import load_dataset
# Example for loading the TREC dataset
dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec")
Available Datasets
Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_LanguageAutoSchemaKGauto_predictive_maintenance_dataAutomated-Personality-PredictionSource:
The dataset is titled PANDORA and is retrieved from the https://psy.takelab.fer.hr/datasets/all/pandora/. the PANDORA dataset is the only dataset that contains personality-relevant information for multiple personality models. It consists of Reddit comments with their corresponding scores for the Big Five Traits, MBTI values and the Enneagrams for more than 10k users.
This Dataset:
This dataset is a subset of Reddit comments from PANDORA focused only on the Big Five Traits. The… See the full description on the dataset page: https://huggingface.co/datasets/Fatima0923/Automated-Personality-Prediction.AutoSchemaKG
AutoSchemaKG
After downloading all the files, please run the following script to merge the files ending with part_* into a single file.
python3 merge_files.py
us-auto-plant-automotive-layoffs-warn-act-notices-daily
US automotive and auto-plant layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-22. 1,315 layoff and closure notices filed by
vehicle assembly plants, auto-parts and powertrain suppliers, tire and axle makers, dealerships and dealer groups, collision and body shops, car washes and vehicle-rental fleets with US state labor departments — 184,299 workers,
665 employers, 39 states, 1989–2027.
276 of the notices (21.0%) were recorded by the state as a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-auto-plant-automotive-layoffs-warn-act-notices-daily.Automated-Essay-Scoring-2.0autotest-sea-ev-dataset
Southeast Asia EV Price, Battery and Range Snapshot
Versioned 4 September 2026 snapshot of battery-electric vehicle variants represented in AutoTest Asia's Thailand and Vietnam datasets.
Files
southeast_asia_ev_snapshot_2026-09-04.csv: harmonized Thailand and Vietnam table, 195 rows.
thailand_ev_snapshot_2026-09-04.csv: 117 Thailand variant rows.
vietnam_ev_snapshot_2026-09-04.csv: 78 Vietnam source rows representing 77 normalized model-variant combinations.… See the full description on the dataset page: https://huggingface.co/datasets/zuozhu625/autotest-sea-ev-dataset.autonlp-data-Ita-Summarizationautoregressive-paraphrase-dataset
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset.autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1What this dataset tests
Whether a system can score traffic-flow coherence
before and after an ego action.
This is not collision detection.
It measures systemic stability.
Required outputs
pre_action_coherence_score
post_action_coherence_score
coherence_delta
shockwave_generation_flag
braking_propagation_depth
systemic_risk_score
Scoring conventions
coherence scores range 0 to 1
coherence_delta may be negative or positive
shockwave flag is 0 or 1
braking propagation depth… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1.deindexing-automation-benchmarks
Deindexing Automation Engine Benchmarks
Benchmark dataset of 20 deindexing automation cases with individual scores for deindex strength, removal rate, review issue handling, reputation health, platform coverage, and workflow efficiency across major removal types and industries.
Built by Deindexing.Services.
Dataset Description
This dataset contains benchmark data for the Deindexing Automation Engine — an automation engine for managing search deindexing, content… See the full description on the dataset page: https://huggingface.co/datasets/deindexing-services/deindexing-automation-benchmarks.Dataset_Automatic_Essay_Scoring_Essay-EssayScore_and_24_textual_featuresautotrain-data-nl-to-sqlautocomplete-search-datasetautotrain-data-numai2Gujarati-Grammarly-DatasetsThis is the collection of datasets used for the creation of our Gujarati Grammarly. It consists of correct-incorrect sentence pairs and correct incorrect spelling pairs datasets.
The sentence pairs are provided in various sizes for ease of prototyping and scaling.
autonomous-driving-social-coherence-field-mapping-v0.1What this dataset tests
Whether a system can score
the coherence of a multi-agent intention field.
This is not collision prediction.
It is social alignment measurement.
Required outputs
dominant_scene_intention
coherence_score
tension_index
conflict_pairs
cooperative_clusters
right_of_way_clarity
Scoring conventions
coherence and tension range 0 to 1
right_of_way_clarity is low, medium, or high
conflict_pairs names agent pairs likely to contest the same space… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-social-coherence-field-mapping-v0.1.icml26-repro-nonlinear-autoencoder-pca
Nonlinear autoencoder / PCA reproduction
This bundle reproduces the paper with the authors' pinned release
(SPOC-group/advantage_nonlinearity@4378017) plus an independent direct
population-gradient-flow audit.
run_official_amp_campaign.py: 96 finite-dimensional runs of the released
two-spike AMP at d=256/512 and eight sample ratios.
run_official_ae_campaign.py: 48 full-batch Adam runs of the released tied,
one-neuron ReLU/ELU autoencoder at d=512/1024, with exact PCA and 20,000… See the full description on the dataset page: https://huggingface.co/datasets/SabaPivot/icml26-repro-nonlinear-autoencoder-pca.Munyarwanda-AI-AutoTrain
Munyarwanda AI - AutoTrain dataset
AutoTrain-ready version of arcange9/Munyarwanda-AI-Dataset v0.2.
Every example is pre-formatted in Qwen chat template as a single text column (5,542 train / 22 validation rows).
Intended recipe (Hugging Face AutoTrain, base model Qwen/Qwen3-0.6B):
LLM task, causal LM
text column: text
LoRA/PEFT + int4 quantization to fit free-tier GPUs
