datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning.
The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering.
The dataset is maintained by with requests to the ArXiv API.
The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.TransitLM
TransitLM: Dataset Release & Evaluation Protocol
Dataset Description
TransitLM is a dataset for public transit route planning in Chinese urban environments, designed to support training and evaluation of language models that generate structured transit routes from origin-destination information. The full dataset covers four cities: Beijing, Shanghai, Shenzhen, and Chengdu, and includes coordinates, station sequences, transfer structure, line information, and route… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/TransitLM.naver-news-summarization-ko
Naver-News-KO: A Korean News Summarization Dataset
A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from
Naver News over a ten-day window in July 2022. It was originally built for a
Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023.
A technical report documenting the collection protocol, corpus statistics, contamination analysis, and
reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.SCASRec
SCASRec: A Self-Correcting and Auto-Stopping Model for Generative Route List Recommendation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Route Features
Used to describe each route, including static features, dynamic features, and trajectory statistical features
N * 62
The estimated time of arrival for the routeThe total distance length of the… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/SCASRec.NACA_4_Digit_for_ML
NACA 4-Digit Airfoil CFD Dataset
Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions.
Dataset Summary
~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles
AoA range: −5° to +5°
Reynolds number range: 100,000 – 500,000
129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.ML-Proto-Datasetmlsum-it
Dataset Card for mlsum-it
Dataset Summary
The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo.
More informations on the official dataset page HuggingFace page.
There are two features:
source: Input news article.
target: Summary of the article.
Supported Tasks and Leaderboards
abstractive-summarization, summarization
Languages
The text in… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/mlsum-it.GenMRP
GenMRP: A Generative Multi-Route Planning Framework for Efficient and Personalized Real-Time Industrial Navigation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Link Features
Includes the road segment attributes
K * 2 * N
Link lengthLink Lane width
Frequency Features
Logs the user's travel history within the past three months
K * 2 * 10 * 7
Delta… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/GenMRP.CCN
Towards Full Candidate Interaction: A Comprehensive Comparison Network for Better Route Recommendation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Route Features
Used to describe each route, including static features, dynamic features, and trajectory statistical features
N * 62
The estimated time of arrival for the routeThe total distance length… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/CCN.MobilityBench
Note: This work is currently under review. The full dataset will be released progressively.
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
Paper | GitHub
MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/MobilityBench.faang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.telco_customer_churnMLMA_hate_speech
Disclaimer
This is a hate speech dataset (in Arabic, French, and English).
Offensive content that does not reflect the opinions of the authors.
Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis)
For more details about our dataset, please check our paper:
@inproceedings{ousidhoum-etal-multilingual-hate-speech-2019,
title = "Multilingual and Multi-Aspect Hate Speech Analysis",
author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.genz-slang-dataset
Dataset Details
This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations.
The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MLBtrio/genz-slang-dataset.pubmed-classification-20k
ml4pubmed/pubmed-classification-20k
20k subset of pubmed text classification from course
R1-Onevision-Bench
R1-Onevision-Bench
[📂 GitHub][📝 Paper]
[🤗 HF Dataset] [🤗 HF Model] [🤗 HF Demo]
Dataset Overview
R1-Onevision-Bench comprises 38 subcategories organized into 5 major domains, including Math, Biology, Chemistry, Physics, Deducation. Additionally, the tasks are categorized into five levels of difficulty, ranging from ‘Junior High School’ to ‘Social Test’ challenges, ensuring a comprehensive evaluation of model capabilities across varying complexities.… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision-Bench.malayalam_news_classificationgerman-credit-risk_credit-scoring_mlp
🏦 German Credit Risk - Dataset para MLP
Este dataset es parte del curso de Deep Learning impartido en el canal de YouTube de inGeniia. Se utiliza para demostrar la implementación de un Perceptrón Multicapa (MLP) para tareas de clasificación binaria (riesgo crediticio).
Descripción del Proyecto
El objetivo de este dataset es predecir si un cliente representa un buen o mal riesgo crediticio basándose en una serie de atributos financieros y personales.
Problema:… See the full description on the dataset page: https://huggingface.co/datasets/inGeniia/german-credit-risk_credit-scoring_mlp.mlb-statcast-batterscoconut-chembl34-selfies-mlm
Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked)
This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks.
The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.ML_for_TwoSampleTesting
Machine Learning for Two-Sample Testing under Right-Censored Data: A Simulation Study
Petr PHILONENKO, Ph.D. in Computer Science;
Sergey POSTOVALOV, D.Sc. in Computer Science.
The paper can be downloaded here.
About
This dataset is a supplement to the github repositiry and paper addressed to solve the two-sample problem under right-censored observations using Machine Learning.
The problem statement can be formualted as H0: S1(t)=S2(t) versus H: S1(t)≠S_2(t) where S1(t)… See the full description on the dataset page: https://huggingface.co/datasets/pfilonenko/ML_for_TwoSampleTesting.ML-Python-Code-Smellsml-phonetic-lexicon
Malayalam Phonetic Lexicon
This dataset contains words in Malayalam script and their pronunciation in International Phonetic Alphabet (IPA)
The words in the lexicon are sourced from
The most frequest 100 thousand words from Indic NLP corpus
Curated collection of word categories from Mlmorph project
This pronunciations are created using Mlphon python Library.
Applications
Ready to use pronunciation lexicons for ASR and TTS
To train datadriven grapheme to phoneme… See the full description on the dataset page: https://huggingface.co/datasets/smcproject/ml-phonetic-lexicon.tox21ASAP_OpenADMET_MLM
ASAP-OpenADMET MLM
ASAP_OpenADMET_MLM dataset from the ASAP Discovery-OpenADMET Antiviral Drug Discovery Challenge [1] [2] [3]. It is intended to be used through
scikit-fingerprints library.
The task is to predict MLM (mouse liver microsomal intrinsic clearance in uL/min/mg) of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
425
Recommended splittime
Recommended metric
MAE
References
[1]
ASAP Discovery
"ASAP… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ASAP_OpenADMET_MLM.mldataML-KEM-SideChannel-Traces
ML-KEM Side Channel Traces
Dataset Description
This dataset contains power traces captured from the Post-Quantum Cryptography (PQC) ML-KEM implementation of the PQM4[1] library (commit: a24bb4b), running on an STM32 Nucleo-L4R5ZI development board equipped with an ARM Cortex-M4 processor.
The traces were collected using a Rohde & Schwarz RTC1002 100 MHz digital oscilloscope. The purpose of this dataset is to evaluate the ML-KEM implementation for side-channel… See the full description on the dataset page: https://huggingface.co/datasets/ai-eldorado/ML-KEM-SideChannel-Traces.ml-word-frequencymlcd-mteb-cifar-eval
MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation
Evaluation results accompanying the MTEB integration of two MLCD image encoders
(PR #5406, resolving
issue #2571).
Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the
reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100
image-classification tasks alongside size-matched OpenAI CLIP baselines.
What was measured
Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.
