CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CShorten /ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning. The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering. The dataset is maintained by with requests to the ArXiv API. The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.tabular100K<n<1M72 likes4k downloads4y agoHugging Face02blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes991 downloads2y agoHugging Face03GD-ML /TransitLM TransitLM: Dataset Release & Evaluation Protocol Dataset Description TransitLM is a dataset for public transit route planning in Chinese urban environments, designed to support training and evaluation of language models that generate structured transit routes from origin-destination information. The full dataset covers four cities: Beijing, Shanghai, Shenzhen, and Chengdu, and includes coordinates, station sequences, transfer structure, line information, and route… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/TransitLM.tabulartext-generation100K<n<1M82 likes977 downloads4mo agoHugging Face04daekeun-ml /naver-news-summarization-ko Naver-News-KO: A Korean News Summarization Dataset A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from Naver News over a ten-day window in July 2022. It was originally built for a Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023. A technical report documenting the collection protocol, corpus statistics, contamination analysis, and reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.textsummarization10K<n<100K65 likes635 downloads2mo agoHugging Face05GD-ML /SCASRec SCASRec: A Self-Correcting and Auto-Stopping Model for Generative Route List Recommendation This is the dataset for our paper. The following table contains the feature dimensions and key features of our dataset. Feature Type Interpretation Shape Some Key Features Route Features Used to describe each route, including static features, dynamic features, and trajectory statistical features N * 62 The estimated time of arrival for the routeThe total distance length of the… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/SCASRec.texttabular-classification100K<n<1M30 likes597 downloads7mo agoHugging Face06Kokoslocke /NACA_4_Digit_for_ML NACA 4-Digit Airfoil CFD Dataset Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions. Dataset Summary ~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles AoA range: −5° to +5° Reynolds number range: 100,000 – 500,000 129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.tabularothern<1K0 likes487 downloads3mo agoHugging Face07gle3D /ML-Proto-Dataset3dn<1K1 likes484 downloads10mo agoHugging Face08ARTeLab /mlsum-it Dataset Card for mlsum-it Dataset Summary The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo. More informations on the official dataset page HuggingFace page. There are two features: source: Input news article. target: Summary of the article. Supported Tasks and Leaderboards abstractive-summarization, summarization Languages The text in… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/mlsum-it.textsummarization10K<n<100K2 likes343 downloads4y agoHugging Face09GD-ML /GenMRP GenMRP: A Generative Multi-Route Planning Framework for Efficient and Personalized Real-Time Industrial Navigation This is the dataset for our paper. The following table contains the feature dimensions and key features of our dataset. Feature Type Interpretation Shape Some Key Features Link Features Includes the road segment attributes K * 2 * N Link lengthLink Lane width Frequency Features Logs the user's travel history within the past three months K * 2 * 10 * 7 Delta… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/GenMRP.texttabular-classification100K<n<1M43 likes342 downloads7mo agoHugging Face10GD-ML /CCN Towards Full Candidate Interaction: A Comprehensive Comparison Network for Better Route Recommendation This is the dataset for our paper. The following table contains the feature dimensions and key features of our dataset. Feature Type Interpretation Shape Some Key Features Route Features Used to describe each route, including static features, dynamic features, and trajectory statistical features N * 62 The estimated time of arrival for the routeThe total distance length… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/CCN.texttabular-classification100K<n<1M37 likes311 downloads7mo agoHugging Face11GD-ML /MobilityBench Note: This work is currently under review. The full dataset will be released progressively. MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios Paper | GitHub MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/MobilityBench.tabularquestion-answering10K<n<100K18 likes297 downloads7mo agoHugging Face12ML-Owl /faang-engineered-time-series-features-2013-2025 FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025) Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!) DOCUMENT NAVIGATION GUIDE (ToC) 1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset 3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.imagetabular-classification10K<n<100K2 likes292 downloads6mo agoHugging Face13MLLab-TS /telco_customer_churntabular1K<n<10K0 likes261 downloads7mo agoHugging Face14nedjmaou /MLMA_hate_speech Disclaimer This is a hate speech dataset (in Arabic, French, and English). Offensive content that does not reflect the opinions of the authors. Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis) For more details about our dataset, please check our paper: @inproceedings{ousidhoum-etal-multilingual-hate-speech-2019, title = "Multilingual and Multi-Aspect Hate Speech Analysis", author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.text10K<n<100K5 likes220 downloads2y agoHugging Face15MLBtrio /genz-slang-dataset Dataset Details This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations. The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MLBtrio/genz-slang-dataset.texttext-generation1K<n<10K52 likes151 downloads2y agoHugging Face16ml4pubmed /pubmed-classification-20k ml4pubmed/pubmed-classification-20k 20k subset of pubmed text classification from course texttext-classification100K<n<1M1 likes118 downloads4y agoHugging Face17Fancy-MLLM /R1-Onevision-Bench R1-Onevision-Bench [📂 GitHub][📝 Paper] [🤗 HF Dataset] [🤗 HF Model] [🤗 HF Demo] Dataset Overview R1-Onevision-Bench comprises 38 subcategories organized into 5 major domains, including Math, Biology, Chemistry, Physics, Deducation. Additionally, the tasks are categorized into five levels of difficulty, ranging from ‘Junior High School’ to ‘Social Test’ challenges, ensuring a comprehensive evaluation of model capabilities across varying complexities.… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision-Bench.textquestion-answeringn<1K3 likes117 downloads2y agoHugging Face18mlexplorer008 /malayalam_news_classificationtext1K<n<10K0 likes89 downloads2y agoHugging Face19inGeniia /german-credit-risk_credit-scoring_mlp 🏦 German Credit Risk - Dataset para MLP Este dataset es parte del curso de Deep Learning impartido en el canal de YouTube de inGeniia. Se utiliza para demostrar la implementación de un Perceptrón Multicapa (MLP) para tareas de clasificación binaria (riesgo crediticio). Descripción del Proyecto El objetivo de este dataset es predecir si un cliente representa un buen o mal riesgo crediticio basándose en una serie de atributos financieros y personales. Problema:… See the full description on the dataset page: https://huggingface.co/datasets/inGeniia/german-credit-risk_credit-scoring_mlp.tabulartabular-classification1K<n<10K3 likes88 downloads10mo agoHugging Face20michaelmallari /mlb-statcast-batterstabular1K<n<10K0 likes86 downloads3y agoHugging Face21gbyuvd /coconut-chembl34-selfies-mlm Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked) This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks. The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.text100K<n<1M0 likes83 downloads1y agoHugging Face22pfilonenko /ML_for_TwoSampleTesting Machine Learning for Two-Sample Testing under Right-Censored Data: A Simulation Study Petr PHILONENKO, Ph.D. in Computer Science; Sergey POSTOVALOV, D.Sc. in Computer Science. The paper can be downloaded here. About This dataset is a supplement to the github repositiry and paper addressed to solve the two-sample problem under right-censored observations using Machine Learning. The problem statement can be formualted as H0: S1(t)=S2(t) versus H: S1(t)≠S_2(t) where S1(t)… See the full description on the dataset page: https://huggingface.co/datasets/pfilonenko/ML_for_TwoSampleTesting.tabular10M<n<100M0 likes81 downloads2y agoHugging Face23jonaskoenig /ML-Python-Code-Smellstexttext-classificationn<1K1 likes80 downloads2y agoHugging Face24smcproject /ml-phonetic-lexicon Malayalam Phonetic Lexicon This dataset contains words in Malayalam script and their pronunciation in International Phonetic Alphabet (IPA) The words in the lexicon are sourced from The most frequest 100 thousand words from Indic NLP corpus Curated collection of word categories from Mlmorph project This pronunciations are created using Mlphon python Library. Applications Ready to use pronunciation lexicons for ASR and TTS To train datadriven grapheme to phoneme… See the full description on the dataset page: https://huggingface.co/datasets/smcproject/ml-phonetic-lexicon.text100K<n<1M1 likes79 downloads3y agoHugging Face25ml-jku /tox21tabular10K<n<100K2 likes78 downloads11mo agoHugging Face26scikit-fingerprints /ASAP_OpenADMET_MLM ASAP-OpenADMET MLM ASAP_OpenADMET_MLM dataset from the ASAP Discovery-OpenADMET Antiviral Drug Discovery Challenge [1] [2] [3]. It is intended to be used through scikit-fingerprints library. The task is to predict MLM (mouse liver microsomal intrinsic clearance in uL/min/mg) of molecules. Characteristic Description Tasks 1 Task type regression Total samples 425 Recommended splittime Recommended metric MAE References [1] ASAP Discovery "ASAP… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ASAP_OpenADMET_MLM.texttabular-regressionn<1K0 likes76 downloads6mo agoHugging Face27cosmicthrillseeking /mldatatext100K<n<1M0 likes69 downloads1y agoHugging Face28ai-eldorado /ML-KEM-SideChannel-Traces ML-KEM Side Channel Traces Dataset Description This dataset contains power traces captured from the Post-Quantum Cryptography (PQC) ML-KEM implementation of the PQM4[1] library (commit: a24bb4b), running on an STM32 Nucleo-L4R5ZI development board equipped with an ARM Cortex-M4 processor. The traces were collected using a Rohde & Schwarz RTC1002 100 MHz digital oscilloscope. The purpose of this dataset is to evaluate the ML-KEM implementation for side-channel… See the full description on the dataset page: https://huggingface.co/datasets/ai-eldorado/ML-KEM-SideChannel-Traces.textother100K<n<1M1 likes64 downloads21d agoHugging Face29kavyamanohar /ml-word-frequencytext1M<n<10M0 likes63 downloads2y agoHugging Face30b4ph /mlcd-mteb-cifar-eval MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation Evaluation results accompanying the MTEB integration of two MLCD image encoders (PR #5406, resolving issue #2571). Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100 image-classification tasks alongside size-matched OpenAI CLIP baselines. What was measured Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.tabularimage-classificationn<1K0 likes60 downloads16d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.