CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes13k downloads2y agoHugging Face02tattabio /ec_classificationtextn<1K0 likes4.7k downloads2y agoHugging Face03amir-kazemi /aidovecl-vehicle-detection-classification-localization AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models. Citation Notice Please ensure that all publications and presentations using this data reference the following paper: Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.imageobject-detection1K<n<10K0 likes4.1k downloads5mo agoHugging Face04jackhhao /jailbreak-classification Jailbreak Classification Dataset Summary Dataset used to classify prompts as jailbreak vs. benign. Dataset Structure Data Fields prompt: an LLM prompt type: classification label, either jailbreak or benign Dataset Creation Curation Rationale Created to help detect & prevent harmful jailbreak prompts when users interact with LLMs. Source Data Jailbreak prompts sourced from: https://github.com/verazuo/jailbreak_llms… See the full description on the dataset page: https://huggingface.co/datasets/jackhhao/jailbreak-classification.texttext-classification1K<n<10K84 likes3.5k downloads3y agoHugging Face05torchsight /cybersecurity-classification-benchmark TorchSight Cybersecurity Classification Benchmark A two-tier benchmark dataset for evaluating cybersecurity document classifiers, released with the TorchSight system. Used in: Dobrovolskyi, I. Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System. Journal of Information Security and Applications, 2026. Canonical per-model numbers live in BENCHMARK_NUMBERS.md, auto-generated from the per-prediction result JSONs… See the full description on the dataset page: https://huggingface.co/datasets/torchsight/cybersecurity-classification-benchmark.texttext-classification1K<n<10K1 likes2k downloads5mo agoHugging Face06SetFit /ade_corpus_v2_classification ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for classification if a sentence is ADE-related (True) or not (False). Train size: 17,637 Test size: 5,879 Source dataset Paper text10K<n<100K6 likes1.9k downloads4y agoHugging Face07mhurhangee /cpc-classificationstext100K<n<1M0 likes1.9k downloads1y agoHugging Face08ccdv /arxiv-classificationArxiv Classification: a classification of Arxiv Papers (11 classes). This dataset is intended for long context classification (documents have all > 4k tokens). Copied from "Long Document Classification From Local Word Glimpses via Recurrent Attention Learning" @ARTICLE{8675939, author={He, Jun and Wang, Liqun and Liu, Liu and Feng, Jiao and Wu, Hao}, journal={IEEE Access}, title={Long Document Classification From Local Word Glimpses via Recurrent Attention Learning}, year={2019}… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/arxiv-classification.texttext-classification10K<n<100K27 likes1.9k downloads2y agoHugging Face09meriemm6 /commit-classification-dataset Commit Classification Dataset This dataset is designed for multi-label classification of Git commit messages into predefined categories. Dataset Summary This dataset contains: Training data: Commit messages and their corresponding labels for training the model. Validation data: A separate set of messages for tuning and evaluation. Testing data: Unlabeled commit messages for testing the model’s performance. The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.texttext-classification1K<n<10K0 likes1.9k downloads2y agoHugging Face10Lots-of-LoRAs /task903_deceptive_opinion_spam_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task903_deceptive_opinion_spam_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task903_deceptive_opinion_spam_classification.texttext-generation1K<n<10K0 likes1.5k downloads2y agoHugging Face11rsh-raj /commit-classification-17ktext10K<n<100K0 likes1.5k downloads2y agoHugging Face12Karavet /ILUR-news-text-classification-corpus News Texts Dataset We release a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens). For more details, refer to the paper. texttext-classification100K<n<1M3 likes1.4k downloads4y agoHugging Face13Lots-of-LoRAs /task902_deceptive_opinion_spam_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task902_deceptive_opinion_spam_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task902_deceptive_opinion_spam_classification.texttext-generation1K<n<10K0 likes1.4k downloads2y agoHugging Face14vnahata /AfriMCQA-category-classification Afri-MCQA cross-modal cultural category classification (MTEB) Classify the cultural category of an entry from its photograph and the question about it spoken by a native speaker, across 16 African languages. Labels index this list: geography, building, and landmarks public figure and pop culture cooking and food objects, materials, clothing tranditions, art, and history brands, products, and companies plants and animals people, and everyday life vehicles and transportation… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-category-classification.audioaudio-classification1K<n<10K0 likes1.2k downloads22d agoHugging Face15Biomedical-TeMU /ProfNER_corpus_classificationtext10K<n<100K2 likes1.1k downloads5y agoHugging Face16dlab-spp /safety-classifications Safety Annotations for dolma3_mix Safety score annotations for a 20K-shard subset of allenai/dolma3_mix-6T using locuslab/safety-classifier_gte-large-en-v1.5. Schema Column Type Description id string Row identifier (matches source dataset) safety_score int8 Argmax safety class (0-5) safety_probs list[float32] Full 6-class probability distribution Safety scale Score Label Count Percentage 0 safe 302,972,734 77.39% 1… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/safety-classifications.texttext-classification100M<n<1B1 likes999 downloads1mo agoHugging Face17ccdv /patent-classificationPatent Classification: a classification of Patents and abstracts (9 classes). This dataset is intended for long context classification (non abstract documents are longer that 512 tokens). Data are sampled from "BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization." by Eva Sharma, Chen Li and Lu Wang See: https://aclanthology.org/P19-1212.pdf See: https://evasharma.github.io/bigpatent/ It contains 9 unbalanced classes, 35k Patents and abstracts divided into 3 splits:… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/patent-classification.texttext-classification10K<n<100K30 likes992 downloads2y agoHugging Face18Kushal0532 /news-political-bias-classification-datasetDataset actually from kaggle. Couldn't find it here so I uploaded it. text10K<n<100K0 likes937 downloads11mo agoHugging Face19ClaudiaRichard /mbti_classification_dataset_fullPoststabular1K<n<10K1 likes934 downloads3y agoHugging Face20ManuD /dfl_classification_512text100K<n<1M1 likes921 downloads4y agoHugging Face21limsc /fr-nfr-classificationtextn<1K2 likes877 downloads4y agoHugging Face22mteb /Vehicle_sounds_classification_datasetaudio1K<n<10K1 likes830 downloads8mo agoHugging Face23CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes788 downloads5mo agoHugging Face24aliencaocao /multimodal_meme_classification_singapore Dataset Card for Offensive Memes in Singapore Context Dataset Details Dataset Description This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards. Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.imagetext-generation100K<n<1M1 likes705 downloads2y agoHugging Face25ibm-nasa-geospatial /multi-temporal-crop-classification Dataset Card for Multi-Temporal Crop Classification Dataset Summary This dataset contains temporal Harmonized Landsat-Sentinel imagery of diverse land cover and crop type classes across the Contiguous United States for the year 2022. The target labels are derived from USDA's Crop Data Layer (CDL). It's primary purpose is for training segmentation geospatial machine learning models. Dataset Structure TIFF Files Each tiff file covers a… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/multi-temporal-crop-classification.text1K<n<10K28 likes692 downloads2y agoHugging Face26nickmuchi /financial-classification Dataset Creation This dataset combines financial phrasebank dataset and a financial text dataset from Kaggle. Given the financial phrasebank dataset does not have a validation split, I thought this might help to validate finance models and also capture the impact of COVID on financial earnings with the more recent Kaggle dataset. texttext-classification1K<n<10K20 likes661 downloads4y agoHugging Face27keremberke /chest-xray-classification Dataset Labels ['NORMAL', 'PNEUMONIA'] Number of Images {'train': 4077, 'test': 582, 'valid': 1165} How to Use Install datasets: pip install datasets Load the dataset: from datasets import load_dataset ds = load_dataset("keremberke/chest-xray-classification", name="full") example = ds['train'][0] Roboflow Dataset Page https://universe.roboflow.com/mohamed-traore-2ekkp/chest-x-rays-qjmia/dataset/2 Citation… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/chest-xray-classification.imageimage-classification1K<n<10K28 likes648 downloads4y agoHugging Face28mteb /multilingual-sentiment-classification MultilingualSentimentClassification An MTEB dataset Massive Text Embedding Benchmark Sentiment classification dataset with binary (positive vs negative sentiment) labels. Includes 30 languages and dialects. Task category t2c DomainsReviews, Written Reference https://huggingface.co/datasets/mteb/multilingual-sentiment-classification How to evaluate on this task You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-sentiment-classification.texttext-classification100K<n<1M1 likes624 downloads1y agoHugging Face29mesolitica /Zeroshot-Audio-Classification-Instructions Zeroshot-Audio-Classification-Instructions Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label, VGGSound FSD50k Nonspeech7k urbansound8K VocalSound Emotion Gender ESD Emotion Age Language TAU Urban Acoustic Scenes 2022 CochlScene BirdCLEF_2021 EmoBox AudioSet We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.audio1M<n<10M4 likes615 downloads1y agoHugging Face30seanswyi /sms-spam-classificationtext1K<n<10K0 likes572 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.