CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Perle-ai /multimodal-ct-radiology-reports Perle AI Multi-phase CECT and CT with Radiology Reports Summary A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering. The release has three configurations: Config Modality Subjects Pairing cect_3phase 3-phase contrast-enhanced abdominal CT (DICOM) 5 per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.tabularimage-classificationn<1K4 likes10k downloads5mo agoHugging Face02jumelet /multiblimp MultiBLiMP MultiBLiMP is a massively Multilingual Benchmark for Linguistic Minimal Pairs. The dataset is composed of synthetic pairs generated using Universal Dependencies and UniMorph. The paper can be found here. We split the data set by language: each language consists of a single .tsv file. The rows contain many attributes for a particular pair, most important are the sen and wrong_sen fields, which we use for evaluating the language models. Using MultiBLiMP To… See the full description on the dataset page: https://huggingface.co/datasets/jumelet/multiblimp.tabular100K<n<1M17 likes7k downloads1y agoHugging Face03OptimusePrime /hle-multimodaltextn<1K1 likes1.7k downloads1y agoHugging Face04facebook /Multi-IF Dataset Summary We introduce Multi-IF, a new benchmark designed to assess LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF, which utilizes a hybrid framework combining LLM and human annotators, expands upon the IFEval by incorporating multi-turn sequences and translating the English prompts into another 7 languages, resulting in a dataset of 4501 multilingual conversations, where each has three turns. Our evaluation of 14 state-of-the-art LLMs on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Multi-IF.tabular1K<n<10K40 likes1.6k downloads2y agoHugging Face05DAMO-NLP-SG /MultiJail Multilingual Jailbreak Challenges in Large Language Models This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models". [Github repo] Annotation Statistics We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below: High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi) Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.textn<1K12 likes1k downloads3y agoHugging Face06Liavan /Traditional-Chinese-Medicine-Multiple_choice_question Discription This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.textquestion-answering1K<n<10K4 likes816 downloads2y agoHugging Face07dbarbedillo /SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset Collection of Multilingual SMS messages tagged as spam or legitimate About Dataset Context The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French. The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dbarbedillo/SMS_Spam_Multilingual_Collection_Dataset.texttext-classification1K<n<10K14 likes758 downloads4y agoHugging Face08AMAImedia /multidomain-kazakh-dataset ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/multidomain-kazakh-dataset.text10M<n<100M0 likes585 downloads9d agoHugging Face09APProjects /us-multi-state-employer-layoffs-warn-notices-by-state Which US employers are laying off in more than one state? Every state publishes its own WARN Act layoff notices, and every state's list stops at its border. The employer that filed in Texas on Monday and Ohio on Wednesday appears as two unrelated rows on two unrelated portals. This dataset is the merge: the free-window notices from 48 states, grouped by employer, kept to employers that filed in two or more states, rebuilt every day. The headline (free window, 1988 -… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-multi-state-employer-layoffs-warn-notices-by-state.tabulartabular-classification10K<n<100K0 likes584 downloads2d agoHugging Face10flax-community /conceptual-12m-mbart-50-multilingualimage10M<n<100M2 likes521 downloads5y agoHugging Face11kevykibbz /ecommerce-behavior-data-from-multi-category-store_oct-nov_2019 eCommerce Behavior Data from Multi-Category Store About the Dataset This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products. Dataset Overview Time Frame: October 2019 - April 2020 Total Events: 285 million Event Granularity: Each row represents an event associated with a product and a user. Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.tabular100M<n<1B4 likes499 downloads2y agoHugging Face12Sp1786 /multiclass-sentiment-analysis-dataset Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.tabulartext-classification10K<n<100K29 likes451 downloads3y agoHugging Face13KumarSahil299885 /SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset Collection of Multilingual SMS messages tagged as spam or legitimate About Dataset Context The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French. The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/KumarSahil299885/SMS_Spam_Multilingual_Collection_Dataset.texttext-classification1K<n<10K3 likes405 downloads3mo agoHugging Face14Johnson8187 /Chinese_Multi-Emotion_Dialogue_Dataset Chinese_Multi-Emotion_Dialogue_Dataset 📄 Description This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text. Data Sources: Daily Conversations: Captured from natural, informal human conversations. Movie Dialogues: Extracted from diverse Chinese-language movies. AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.texttext-classification1K<n<10K19 likes367 downloads12d agoHugging Face15BothBosu /multi-agent-scam-conversation Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities Dataset Description The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.texttext-classification1K<n<10K10 likes324 downloads2y agoHugging Face16krammnic /hle-multichoiceHumanity Last Exam dataset with extra incorrect answers generated with Qwen3-4B tabulartable-question-answering1K<n<10K0 likes324 downloads1y agoHugging Face17kz-transformers /multidomain-kazakh-dataset Dataset Description Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk Dataset Summary MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains. Supported Tasks 'MLM/CLM': can be used to train a model for casual and masked languange modeling Languages The kk code for Kazakh as generally spoken in the Kazakhstan Data Instances For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.texttext-generation10M<n<100M30 likes307 downloads1y agoHugging Face18ameyhengle /Multilingual-Needle-in-a-Haystack Multilingual Needle in a Haystack (MLNeedle) The MultiLingual Needle-in-a-Haystack (MLNeedle) test is a dataset designed to assess how well Large Language Models (LLMs) find specific information ("needle") within long, multilingual texts ("haystack"). Built on MLQA, it contains over 5,000 extractive question-answer instances across seven languages (English, Arabic, German, Spanish, Hindi, Vietnamese, Simplified Chinese). We systematically vary the "needle's" language and position to… See the full description on the dataset page: https://huggingface.co/datasets/ameyhengle/Multilingual-Needle-in-a-Haystack.text10K<n<100K3 likes296 downloads1y agoHugging Face19mcaleste /sat_multiple_choice_math_may_23This is the set of math SAT questions from the May 2023 SAT, taken from here: https://www.mcelroytutoring.com/lower.php?url=44-official-sat-pdfs-and-82-official-act-pdf-practice-tests-free. Questions that included images were not included but all other math questions, including those that have tables were included. tabularn<1K2 likes281 downloads3y agoHugging Face20neemiasbsilva /multimodal-LLMs-See-Sentiment MLLMsent — datasets and experiment results Every input and every output of "Multimodal LLMs See Sentiment" (arXiv:2508.16873): the image descriptions generated by six multimodal LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete per-fold results of all 141 experiments. Paper: arXiv:2508.16873 Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.texttext-classification10K<n<100K1 likes263 downloads1mo agoHugging Face21Multilingual-Perspectivist-NLU /MultiPICo Dataset Summary MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.tabular10K<n<100K6 likes259 downloads2y agoHugging Face22trucyberlab /multimodal-ICS-provenance ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.tabulargraph-ml100K<n<1M0 likes247 downloads3mo agoHugging Face23Kenpache /multilingual-financial-sentiment Multilingual Financial Sentiment Dataset A curated dataset of 39,829 financial news sentences annotated with sentiment labels (Negative / Neutral / Positive) across 7 languages, collected from 80+ financial news sources worldwide. Dataset Summary Total samples 39,829 Languages 7 (EN, ZH, JA, DE, FR, ES, AR) Labels 3 (negative, neutral, positive) Format CSV Sources 80+ financial news outlets Languages Language Code Samples %… See the full description on the dataset page: https://huggingface.co/datasets/Kenpache/multilingual-financial-sentiment.texttext-classification10K<n<100K0 likes236 downloads6mo agoHugging Face24owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K29 likes222 downloads4y agoHugging Face25alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes220 downloads5mo agoHugging Face26multimodalart /panda-70m Panda 70M dataset by Snap Inc 70M video-caption pairs Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading textimage-to-text100K<n<1M15 likes216 downloads2y agoHugging Face27flax-community /conceptual-12m-multilingual-marianThis dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following: train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each) val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.text10M<n<100M1 likes212 downloads3y agoHugging Face28maximoss /lingnli-multi Dataset Card for Dataset Name Dataset Summary This repository contains a collection of machine translations of LingNLI dataset into 9 different languages (Bulgarian, Finnish, French, Greek, Italian, Korean, Lithuanian, Portuguese, Spanish). The goal is to predict textual entailment (does sentence A imply/contradict/neither sentence B), which is a classification task (given two sentences, predict one of three labels). It is here formatted in the same manner as the… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/lingnli-multi.texttext-classification100K<n<1M1 likes212 downloads2y agoHugging Face29AdityaaXD /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024). 📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K7 likes204 downloads8mo agoHugging Face30NorGLM /NO-Multi-QA-Sum Dataset Card Dataset Summary NO-Multi-QA-Sum is a Norwegian multi-task human annotated dataset. It is a part of NLEBench Norwegian benchmarks, and can be used for evaluation of Machine reading comprehension, document-grounded question answering, abstractive summarization tasks of Language Models. Language The data in NO-Alpaca-Plus are in Norwegian Bokmål. Data Instances For each instance, there is an article string, category, summary string, and a… See the full description on the dataset page: https://huggingface.co/datasets/NorGLM/NO-Multi-QA-Sum.textn<1K1 likes191 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.