CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Perle-ai /multimodal-ct-radiology-reports Perle AI Multi-phase CECT and CT with Radiology Reports Summary A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering. The release has three configurations: Config Modality Subjects Pairing cect_3phase 3-phase contrast-enhanced abdominal CT (DICOM) 5 per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.tabularimage-classificationn<1K4 likes10k downloads5mo agoHugging Face02jumelet /multiblimp MultiBLiMP MultiBLiMP is a massively Multilingual Benchmark for Linguistic Minimal Pairs. The dataset is composed of synthetic pairs generated using Universal Dependencies and UniMorph. The paper can be found here. We split the data set by language: each language consists of a single .tsv file. The rows contain many attributes for a particular pair, most important are the sen and wrong_sen fields, which we use for evaluating the language models. Using MultiBLiMP To… See the full description on the dataset page: https://huggingface.co/datasets/jumelet/multiblimp.tabular100K<n<1M17 likes7k downloads1y agoHugging Face03facebook /Multi-IF Dataset Summary We introduce Multi-IF, a new benchmark designed to assess LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF, which utilizes a hybrid framework combining LLM and human annotators, expands upon the IFEval by incorporating multi-turn sequences and translating the English prompts into another 7 languages, resulting in a dataset of 4501 multilingual conversations, where each has three turns. Our evaluation of 14 state-of-the-art LLMs on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Multi-IF.tabular1K<n<10K40 likes1.6k downloads2y agoHugging Face04APProjects /us-multi-state-employer-layoffs-warn-notices-by-state Which US employers are laying off in more than one state? Every state publishes its own WARN Act layoff notices, and every state's list stops at its border. The employer that filed in Texas on Monday and Ohio on Wednesday appears as two unrelated rows on two unrelated portals. This dataset is the merge: the free-window notices from 48 states, grouped by employer, kept to employers that filed in two or more states, rebuilt every day. The headline (free window, 1988 -… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-multi-state-employer-layoffs-warn-notices-by-state.tabulartabular-classification10K<n<100K0 likes584 downloads2d agoHugging Face05kevykibbz /ecommerce-behavior-data-from-multi-category-store_oct-nov_2019 eCommerce Behavior Data from Multi-Category Store About the Dataset This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products. Dataset Overview Time Frame: October 2019 - April 2020 Total Events: 285 million Event Granularity: Each row represents an event associated with a product and a user. Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.tabular100M<n<1B4 likes499 downloads2y agoHugging Face06Sp1786 /multiclass-sentiment-analysis-dataset Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.tabulartext-classification10K<n<100K29 likes451 downloads3y agoHugging Face07mulab-mir /muchomusic MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio. It includes 1,187 multiple-choice questions validated by human annotators, based on 644 music tracks from two publicly available music datasets. These questions cover a wide variety of genres and assess knowledge and reasoning across several musical concepts and their cultural and functional… See the full description on the dataset page: https://huggingface.co/datasets/mulab-mir/muchomusic.tabular1K<n<10K8 likes401 downloads2y agoHugging Face08krammnic /hle-multichoiceHumanity Last Exam dataset with extra incorrect answers generated with Qwen3-4B tabulartable-question-answering1K<n<10K0 likes324 downloads1y agoHugging Face09mcaleste /sat_multiple_choice_math_may_23This is the set of math SAT questions from the May 2023 SAT, taken from here: https://www.mcelroytutoring.com/lower.php?url=44-official-sat-pdfs-and-82-official-act-pdf-practice-tests-free. Questions that included images were not included but all other math questions, including those that have tables were included. tabularn<1K2 likes281 downloads3y agoHugging Face10Multilingual-Perspectivist-NLU /MultiPICo Dataset Summary MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.tabular10K<n<100K6 likes259 downloads2y agoHugging Face11trucyberlab /multimodal-ICS-provenance ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.tabulargraph-ml100K<n<1M0 likes247 downloads3mo agoHugging Face12owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K29 likes222 downloads4y agoHugging Face13AdityaaXD /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024). 📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K7 likes204 downloads8mo agoHugging Face14FrancophonIA /multilingual-hatespeech-dataset [!NOTE] Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset Description This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate texts also the data from different languages needed to be identified as a corresponding correct language. The following are the languages in the dataset with the numbers corresponding to that language. (1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.tabular100K<n<1M4 likes176 downloads1y agoHugging Face15sanjaydoss /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K10 likes174 downloads23d agoHugging Face16Multilingual-Perspectivist-NLU /EPIC Dataset Card for EPICorpus Dataset Summary EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.tabulartext-classification10K<n<100K2 likes161 downloads2y agoHugging Face17oliviersportsdata /NFL-Multi-Market-Timestamped-Odds-Team-Stats 🏈 NFL Multi-Market — Timestamped Odds & Team Stats (2018–2026) Free sample — 18 games, 26,169 timestamped snapshots, eight Super Bowls. Every NFL archive on the market gives you one closing number per game. The full dataset gives you the entire line history: 2,816,380 timestamped snapshots across 2,242 games, two named books and three markets — pre-match and in-running, with the live score and the game clock on the same row as the price. → Get the full archive — 2,816,380… See the full description on the dataset page: https://huggingface.co/datasets/oliviersportsdata/NFL-Multi-Market-Timestamped-Odds-Team-Stats.tabulartabular-regression10K<n<100K0 likes141 downloads7d agoHugging Face18Multimedia-SMU /seeingculture-benchmarkPaper | Project Page | Leaderboard | Explorer | Code | CMB, the video successor Seeing Culture Benchmark (SCB) Evaluating Visual Reasoning and Grounding in Cultural Context The Seeing Culture Benchmark (SCB) evaluates cultural reasoning in vision-language models in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual options in… See the full description on the dataset page: https://huggingface.co/datasets/Multimedia-SMU/seeingculture-benchmark.imageimage-text-to-text1K<n<10K3 likes138 downloads18h agoHugging Face19AAAIBenchmark /Multi-Opthalingua Cite Accepted to AAAI 2025 (https://openreview.net/group?id=AAAI.org/2025/Conference#tab-recent-activity) Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs: @misc{restrepo2024multiophthalinguamultilingualbenchmarkassessing, title={Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs}, author={David Restrepo and Chenwei Wu and Zhengxu Tang and Zitao Shuai and Thao… See the full description on the dataset page: https://huggingface.co/datasets/AAAIBenchmark/Multi-Opthalingua.tabularquestion-answeringn<1K3 likes123 downloads2y agoHugging Face20gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes115 downloads2y agoHugging Face21Booking-com /multi-destination-trip-dataset Intro Booking.com provides a unique dataset based on millions of real anonymized bookings to encourage the research on sequential recommendation problems. Many travelers go on trips which include more than one destination. Our mission at Booking.com is to make it easier for everyone to experience the world, and we can help to do that by providing real-time recommendations for what their next in-trip destination will be. By making accurate predictions, we help deliver a frictionless… See the full description on the dataset page: https://huggingface.co/datasets/Booking-com/multi-destination-trip-dataset.tabular1M<n<10M6 likes103 downloads2y agoHugging Face22danielelvs /multilingual-islr-mediapipe Multilingual ISLR MediaPipe Landmarks Dataset Description This dataset combines frame-level MediaPipe Holistic landmarks derived from four isolated sign language recognition (ISLR) resources: INCLUDE-50, KSL, MINDS-Libras, and LIBRAS-UFOP. It provides a common tabular schema for research on landmark selection, temporal modeling, signer-independent evaluation, and multilingual transfer learning. The release contains landmarks rather than source RGB videos. Every… See the full description on the dataset page: https://huggingface.co/datasets/danielelvs/multilingual-islr-mediapipe.tabularvideo-classification1K<n<10K0 likes93 downloads9d agoHugging Face23oliviersportsdata /FIFA-World-Cup-Multi-Market-Timestamped-Odds-Match-Stats FIFA World Cup — In-Running Timestamped Odds (free sample) This repository holds a free sample: 10 matches out of 232, drawn across the three editions and both phases of each tournament, including two penalty shootouts. It is published so the schema can be inspected before purchase. Edition Rounds in the sample 2018 Round 1 · Round 3 · Round of 16 (shootout) 2022 Round 1 · Round 2 · Round of 16 (shootout) 2026 Round 1 · Round 3 · Round of 32 (shootout) · Round of… See the full description on the dataset page: https://huggingface.co/datasets/oliviersportsdata/FIFA-World-Cup-Multi-Market-Timestamped-Odds-Match-Stats.tabulartabular-regression10K<n<100K0 likes91 downloads6d agoHugging Face24FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes86 downloads5mo agoHugging Face25maximoss /rte3-multi Dataset Card for multilingual RTE-3 Dataset Summary This repository contains all manually translated versions of RTE-3 dataset, plus the original English one. The languages into which RTE-3 dataset has so far been translated are Italian (2012), German (2013), and French (2023). Unlike in other repositories, both our own French version and the older Italian and German ones are here annotated in 3 classes (entailment, neutral, contradiction), and not in 2 (entailment, not… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/rte3-multi.tabulartext-classification1K<n<10K2 likes85 downloads7mo agoHugging Face26SM-Bello /ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark 🛸 ZeroTwin-UAV-Synthetic Multi-Agent Physics-Informed Degradation Benchmark for Autonomous UAV Swarms ═══════════════════════════════════════════════════════════════════════════════════════ P H I L A B • P E N E L O P E I N C . R E S E A R C H D I V I S I O N ═══════════════════════════════════════════════════════════════════════════════════════ 🏛️ Provenance & Institutional Trademarks This open-source benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark.tabulartime-series-forecasting10K<n<100K0 likes84 downloads1mo agoHugging Face27AdityaaXD /Multi-Model-Trading-Data 📊 Multi-Model Trading Data Bitcoin (BTC-USD) historical price data with technical indicators for ML/DL trading models. 📁 Dataset Files File Description Rows Columns btc_usd_historical.csv Raw OHLCV data ~3,653 5 btc_usd_features.csv Processed with indicators ~3,603 17 📅 Date Range Start: 2015-01-01 End: 2025-01-01 Frequency: Daily 📈 Features in btc_usd_features.csv Raw OHLCV open, high, low, close, volume… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Model-Trading-Data.tabulartabular-classification1K<n<10K0 likes83 downloads8mo agoHugging Face28Steveeeeeeen /multilingual_evalstabularn<1K0 likes82 downloads4mo agoHugging Face29Rapidata /multilingual-llm-jokes-4o-claude-gemini Rapidata Generated Joke Preference Dataset We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'. It took us less than 5 days to get all of the responses. The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.tabular1K<n<10K14 likes78 downloads1y agoHugging Face30Muleex12 /tourism-package-prediction-datatabular1K<n<10K0 likes77 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.