CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rami /multi-label-class-github-issues-text-classification Dataset Card for "multi-label-class-github-issues-text-classification" More Information needed text1K<n<10K2 likes319 downloads4y agoHugging Face02argilla /research_titles_multi-labeltext10K<n<100K0 likes270 downloads4y agoHugging Face03owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K27 likes210 downloads4y agoHugging Face04HC-85 /arxiv-abstract-multilabeltabulartext-classification1M<n<10M0 likes194 downloads2y agoHugging Face05marcelsun /wos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification. The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.texttext-classification100K<n<1M0 likes183 downloads2y agoHugging Face06suwaimyo /gklmip-news-khm-multilabelclassification GKLMIPNews_khm_MultiLabelClassification Deduplicated copy of kornwtp/gklmip-news-khm-multilabelclassification. Splits split rows test 1,398 train 3,959 validation 1,376 text1K<n<10K0 likes143 downloads21d agoHugging Face07suwaimyo /vlsp2018sa-hotel-vie-multilabelclassification VLSP2018SAHotel_vie_MultiLabelClassification Deduplicated copy of kornwtp/vlsp2018sa-hotel-vie-multilabelclassification. Splits split rows test 599 train 2,947 validation 1,985 text1K<n<10K0 likes141 downloads21d agoHugging Face08suwaimyo /hatespeech-ind-multilabelclassification HateSpeech_ind_MultiLabelClassification Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification. Splits split rows train 13,014 text10K<n<100K0 likes135 downloads21d agoHugging Face09suwaimyo /netifier-ind-multilabelclassification Netifier_ind_MultiLabelClassification Deduplicated copy of kornwtp/netifier-ind-multilabelclassification. Splits split rows test 750 train 6,361 text1K<n<10K0 likes127 downloads21d agoHugging Face10IvanFed /russian-toxic-comments-multilabel Russian Toxic Comments Multi-label Dataset Dataset Description Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности. Цель Обучение модели для автоматического обнаружения трех типов токсичного контента: Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.tabulartext-classification100K<n<1M1 likes103 downloads2mo agoHugging Face11suwaimyo /prachathai67k-tha-multilabelclassification Prachathai67k_tha_MultiLabelClassification Deduplicated copy of kornwtp/prachathai67k-tha-multilabelclassification. Splits split rows train 67,488 text10K<n<100K0 likes102 downloads25d agoHugging Face12clip-benchmark /wds_voc2007_multilabelimage1K<n<10K1 likes93 downloads4y agoHugging Face13suwaimyo /casa-ind-multilabelclassification CASA_ind_MultiLabelClassification Deduplicated copy of kornwtp/casa-ind-multilabelclassification. Splits split rows test 180 train 809 validation 90 text1K<n<10K0 likes92 downloads25d agoHugging Face14suwaimyo /dengue-fil-multilabelclassification Dengue_fil_MultiLabelClassification Deduplicated copy of kornwtp/dengue-fil-multilabelclassification. Splits split rows test 494 train 3,924 validation 498 text1K<n<10K0 likes92 downloads25d agoHugging Face15suwaimyo /vlsp2018sa-restaurant-vie-multilabelclassification VLSP2018SARestaurant_vie_MultiLabelClassification Deduplicated copy of kornwtp/vlsp2018sa-restaurant-vie-multilabelclassification. Splits split rows test 499 train 2,958 validation 1,289 text1K<n<10K0 likes91 downloads25d agoHugging Face16kornwtp /prachathai67k-tha-multilabelclassificationref: https://github.com/PyThaiNLP/prachathai-67k text10K<n<100K0 likes90 downloads2y agoHugging Face17FiscaAI /icd10cm-multilabel-prompttext100K<n<1M4 likes86 downloads2y agoHugging Face18taishi-i /awesome-japanese-nlp-multilabel-dataset Dataset overview This is a dataset for Japanese natural language processing with multi-label annotations of research field labels for GitHub repositories in the NLP domain. Please refer to this paper for the specific method of constructing the dataset. It is written in Japanese. Input and Output Input: Information from GitHub repositories (description, README text, PDF text, screenshot images) Output: Multi-label classification of NLP research fields Problem Setting of the… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/awesome-japanese-nlp-multilabel-dataset.texttext-classificationn<1K0 likes85 downloads2y agoHugging Face19suwaimyo /hoasa-ind-multilabelclassification HoASA_ind_MultiLabelClassification Deduplicated copy of kornwtp/hoasa-ind-multilabelclassification. Splits split rows test 286 train 2,267 validation 285 text1K<n<10K0 likes85 downloads25d agoHugging Face20Talip7 /scikit-learn-issues-multilabel 🧩 Scikit-learn GitHub Issues – Multilabel Dataset This dataset contains GitHub issues from the scikit-learn repository, prepared for multilabel NLP tasks such as issue tagging, automated triage, and semantic search. Each row corresponds to one issue-comment context, making the dataset suitable for real-world developer tooling. 📌 Motivation GitHub issues are a critical signal in open-source projects: Bug tracking Feature requests Documentation improvements… See the full description on the dataset page: https://huggingface.co/datasets/Talip7/scikit-learn-issues-multilabel.texttext-classification10K<n<100K0 likes81 downloads9mo agoHugging Face21ibrahimdaud /multi-label-food-recognition Multi-Label Food Recognition Dataset This is a multi-label food recognition dataset generated from single-class food images. Each image contains 2-5 different food items composited together using natural composition methods. Dataset Details Total Images: 13,000 Training Images: 10,400 (80%) Validation Images: 2,600 (20%) Number of Classes: 90 Labels per Image: 2-5 labels Image Format: RGB, 512x512 pixels File Format: Parquet Dataset Structure Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.imageimage-classification10K<n<100K1 likes80 downloads10mo agoHugging Face22suwaimyo /prachathai67k-mya-multilabelclassification Prachathai67k_mya_MultiLabelClassification Deduplicated copy of kornwtp/prachathai67k-mya-multilabelclassification. Splits split rows train 2,153 text1K<n<10K0 likes80 downloads25d agoHugging Face23suwaimyo /truevoice-intent-tha-multilabelclassification TrueVoiceIntent_tha_MultiLabelClassification Deduplicated copy of kornwtp/truevoice-intent-tha-multilabelclassification. Splits split rows train 13,355 text10K<n<100K0 likes79 downloads25d agoHugging Face24acloudfan /toxicity-multi-label-classifier Part of a course titled "Generative AI application design & development" https://genai.acloudfan.com/ Created from a dataset available on Kaggle. https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data tabulartext-classificationn<1K0 likes78 downloads2y agoHugging Face25argilla /synthetic-text-classification-news-multi-label Dataset Card for synthetic-text-classification-news-multi-label This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/synthetic-text-classification-news-multi-label/raw/main/pipeline.yaml" or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news-multi-label.textn<1K6 likes77 downloads2y agoHugging Face26peterbeamish /environment-multi-labels-even Dataset Card for "environment-multi-labels-even" More Information needed tabular10K<n<100K0 likes73 downloads3y agoHugging Face27k3999 /multilabel-imagenet-1k MultiLabel ImageNet-1K Train Annotations with Selected Masks This dataset contains automated multi-label annotations for the ImageNet-1K training split, together with spatial masks for the selected object-level labels. The release is designed to make the annotations easy to inspect and reuse. It does not include the original ImageNet images. Users need access to the ImageNet-1K training images separately; image paths are stored relative to the ImageNet train root, for example:… See the full description on the dataset page: https://huggingface.co/datasets/k3999/multilabel-imagenet-1k.textimage-classification1M<n<10M1 likes66 downloads4mo agoHugging Face28dbrovkin /toxic-russian-comments-multilabel Russian Toxic Comments Multi-Label Balanced Dataset Описание Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками: profanity: наличие нецензурной лексики (мат) threat: наличие угроз illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT) Структура данных Датасет содержит следующие поля: Поле Тип Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.tabulartext-classification100K<n<1M1 likes64 downloads2mo agoHugging Face29jakeazcona /short-text-multi-labeled-emotion-classificationtabular10K<n<100K2 likes61 downloads5y agoHugging Face30kornwtp /vlsp2018sa-restaurant-vie-multilabelclassificationref: https://github.com/vndee/awsome-vietnamese-nlp text1K<n<10K0 likes58 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.