CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /finepdfs_lang_classificationtabular1M<n<10M4 likes18k downloads11mo agoHugging Face02tattabio /ec_classificationtextn<1K0 likes4.7k downloads2y agoHugging Face03ccdv /arxiv-classificationArxiv Classification: a classification of Arxiv Papers (11 classes). This dataset is intended for long context classification (documents have all > 4k tokens). Copied from "Long Document Classification From Local Word Glimpses via Recurrent Attention Learning" @ARTICLE{8675939, author={He, Jun and Wang, Liqun and Liu, Liu and Feng, Jiao and Wu, Hao}, journal={IEEE Access}, title={Long Document Classification From Local Word Glimpses via Recurrent Attention Learning}, year={2019}… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/arxiv-classification.texttext-classification10K<n<100K27 likes1.7k downloads2y agoHugging Face04RoboCOIN /Cobot_Magic_classification_of_tablewaregated Cobot_Magic_classification_of_tableware 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_tableware.tabularrobotics100K<n<1M0 likes1.5k downloads9mo agoHugging Face05Lots-of-LoRAs /task903_deceptive_opinion_spam_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task903_deceptive_opinion_spam_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task903_deceptive_opinion_spam_classification.texttext-generation1K<n<10K0 likes1.5k downloads2y agoHugging Face06Lots-of-LoRAs /task902_deceptive_opinion_spam_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task902_deceptive_opinion_spam_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task902_deceptive_opinion_spam_classification.texttext-generation1K<n<10K0 likes1.4k downloads2y agoHugging Face07RoboCOIN /Cobot_Magic_classification_of_fruits_and_vegetablesgated Cobot_Magic_classification_of_fruits_and_vegetables 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables.tabularrobotics100K<n<1M0 likes1.3k downloads9mo agoHugging Face08vnahata /AfriMCQA-category-classification Afri-MCQA cross-modal cultural category classification (MTEB) Classify the cultural category of an entry from its photograph and the question about it spoken by a native speaker, across 16 African languages. Labels index this list: geography, building, and landmarks public figure and pop culture cooking and food objects, materials, clothing tranditions, art, and history brands, products, and companies plants and animals people, and everyday life vehicles and transportation… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-category-classification.audioaudio-classification1K<n<10K0 likes1.2k downloads21d agoHugging Face09RoboCOIN /Cobot_Magic_classification_of_fruits_and_vegetables_agated Cobot_Magic_classification_of_fruits_and_vegetables_a 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables_a.tabularrobotics100K<n<1M0 likes1k downloads9mo agoHugging Face10ClaudiaRichard /mbti_classification_dataset_fullPoststabular1K<n<10K1 likes924 downloads3y agoHugging Face11ccdv /patent-classificationPatent Classification: a classification of Patents and abstracts (9 classes). This dataset is intended for long context classification (non abstract documents are longer that 512 tokens). Data are sampled from "BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization." by Eva Sharma, Chen Li and Lu Wang See: https://aclanthology.org/P19-1212.pdf See: https://evasharma.github.io/bigpatent/ It contains 9 unbalanced classes, 35k Patents and abstracts divided into 3 splits:… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/patent-classification.texttext-classification10K<n<100K30 likes887 downloads2y agoHugging Face12limsc /fr-nfr-classificationtextn<1K2 likes878 downloads4y agoHugging Face13ManuD /dfl_classification_512text100K<n<1M1 likes847 downloads4y agoHugging Face14mteb /Vehicle_sounds_classification_datasetaudio1K<n<10K1 likes817 downloads8mo agoHugging Face15nickmuchi /financial-classification Dataset Creation This dataset combines financial phrasebank dataset and a financial text dataset from Kaggle. Given the financial phrasebank dataset does not have a validation split, I thought this might help to validate finance models and also capture the impact of COVID on financial earnings with the more recent Kaggle dataset. texttext-classification1K<n<10K20 likes700 downloads4y agoHugging Face16mteb /multilingual-sentiment-classification MultilingualSentimentClassification An MTEB dataset Massive Text Embedding Benchmark Sentiment classification dataset with binary (positive vs negative sentiment) labels. Includes 30 languages and dialects. Task category t2c DomainsReviews, Written Reference https://huggingface.co/datasets/mteb/multilingual-sentiment-classification How to evaluate on this task You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-sentiment-classification.texttext-classification100K<n<1M1 likes662 downloads1y agoHugging Face17aliencaocao /multimodal_meme_classification_singapore Dataset Card for Offensive Memes in Singapore Context Dataset Details Dataset Description This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards. Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.imagetext-generation100K<n<1M1 likes637 downloads2y agoHugging Face18GATE-engine /happy-whale-dolphin-classificationimage10K<n<100K3 likes621 downloads3y agoHugging Face19mesolitica /Zeroshot-Audio-Classification-Instructions Zeroshot-Audio-Classification-Instructions Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label, VGGSound FSD50k Nonspeech7k urbansound8K VocalSound Emotion Gender ESD Emotion Age Language TAU Urban Acoustic Scenes 2022 CochlScene BirdCLEF_2021 EmoBox AudioSet We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.audio1M<n<10M3 likes613 downloads1y agoHugging Face20mteb /multilingual-scala-classification ScalaClassification An MTEB dataset Massive Text Embedding Benchmark ScaLa a linguistic acceptability dataset for the mainland Scandinavian languages automatically constructed from dependency annotations in Universal Dependencies Treebanks. Published as part of 'ScandEval: A Benchmark for Scandinavian Natural Language Processing' Task category t2c Domains Fiction, News, Non-fiction, Blog, Spoken, Web, Written Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-scala-classification.texttext-classification10K<n<100K1 likes579 downloads7mo agoHugging Face21TheFinAI /fiqa-sentiment-classification Dataset Name Dataset Description This dataset is based on the task 1 of the Financial Sentiment Analysis in the Wild (FiQA) challenge. It follows the same settings as described in the paper 'A Baseline for Aspect-Based Sentiment Analysis in Financial Microblogs and News'. The dataset is split into three subsets: train, valid, test with sizes 822, 117, 234 respectively. Dataset Structure _id: ID of the data point sentence: The sentence target: The target of the… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/fiqa-sentiment-classification.text1K<n<10K7 likes534 downloads3y agoHugging Face22Jayabalambika /hagrid-classification-512p-dataset Dataset Card for "hagrid-classification-512p-dataset" More Information needed image100K<n<1M0 likes528 downloads3y agoHugging Face23C-MTEB /TNews-classification Dataset Card for "TNews-classification" More Information needed text10K<n<100K2 likes511 downloads3y agoHugging Face24kenhktsui /code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile. It is intended to be used for training code natural language classifier. texttext-classification1M<n<10M0 likes477 downloads2y agoHugging Face25raghavendrad60 /vqa_plant-disease-classification-merged-datasetimage10K<n<100K1 likes461 downloads2y agoHugging Face26rpmon /fma-genre-classification FMA Genre Classification Dataset The FMA Genre Classification Dataset is a subset of the Free Music Archive (FMA), containing audio samples and genre labels for music classification tasks. This version uses the "small" subset of FMA, which contains 8,000 tracks of 30 seconds each, evenly distributed across 8 genres. Dataset Description Dataset Summary This dataset consists of 8,000 audio tracks from the Free Music Archive (FMA), each 30 seconds in length… See the full description on the dataset page: https://huggingface.co/datasets/rpmon/fma-genre-classification.audio1K<n<10K3 likes459 downloads2y agoHugging Face27mwirth-epo /cpc-classification-data CPC classification datasets These datasets have been used to train the CPC (Cooperative Patent Classification) classification models mentioned in the article Hähnke, V. D., Wéry, A., Wirth, M., & Klenner-Bajaja, A. (2025). Encoder models at the European Patent Office: Pre-training and use cases. World Patent Information, 81, 102360. https://doi.org/10.1016/j.wpi.2025.102360. Columns: publication_number: the patent publication number, the content of the publication can be looked up… See the full description on the dataset page: https://huggingface.co/datasets/mwirth-epo/cpc-classification-data.texttext-classification10M<n<100M1 likes453 downloads1y agoHugging Face28jy13 /bi-so101-fruits-classificationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so101_follower", "total_episodes": 2, "total_frames": 2910, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/bi-so101-fruits-classification.tabularrobotics10K<n<100K0 likes340 downloads1y agoHugging Face29C-MTEB /OnlineShopping-classification Dataset Card for "OnlineShopping-classification" More Information needed text1K<n<10K4 likes336 downloads3y agoHugging Face30luisgasco /profner_classification_master Binary Classification Dataset: Profession Detection in Tweets This dataset is a derived version of the original PROFNER task, adapted for binary text classification. The goal is to determine whether a tweet mentions a profession or not. 🧠 Objective Each example contains: A tweet_id (document identifier) A text field (full tweet content) A label, which has been normalized into two classes: CON_PROFESION: The tweet contains a reference to a profession. SIN_PROFESION: The… See the full description on the dataset page: https://huggingface.co/datasets/luisgasco/profner_classification_master.text1K<n<10K1 likes335 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.