CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K28 likes221 downloads4y agoHugging Face02HC-85 /arxiv-abstract-multilabeltabulartext-classification1M<n<10M0 likes191 downloads2y agoHugging Face03IvanFed /russian-toxic-comments-multilabel Russian Toxic Comments Multi-label Dataset Dataset Description Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности. Цель Обучение модели для автоматического обнаружения трех типов токсичного контента: Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.tabulartext-classification100K<n<1M1 likes108 downloads2mo agoHugging Face04dbrovkin /toxic-russian-comments-multilabel Russian Toxic Comments Multi-Label Balanced Dataset Описание Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками: profanity: наличие нецензурной лексики (мат) threat: наличие угроз illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT) Структура данных Датасет содержит следующие поля: Поле Тип Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.tabulartext-classification100K<n<1M1 likes62 downloads2mo agoHugging Face05puttatidam /hoasa-ind-multilabelclassificationtabularn<1K0 likes43 downloads7d agoHugging Face06puttatidam /dengue-fil-multilabelclassificationtabularn<1K0 likes42 downloads7d agoHugging Face07syke9p3 /multilabel-tagalog-hate-speechtabular1K<n<10K0 likes35 downloads2y agoHugging Face08TatarNLPWorld /tatar-news-analysis-multilabelgated Dataset Card for Tatar News Multilabel Classification Dataset Details Dataset Description The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.tabulartext-classification10K<n<100K0 likes34 downloads25d agoHugging Face09chillies /course-review-multilabel-sentiment-analysistabular1K<n<10K0 likes32 downloads2y agoHugging Face10puttatidam /casa-ind-multilabelclassificationtabularn<1K0 likes27 downloads7d agoHugging Face11maximuspowers /philosophy-schools-multilabel Dataset Card for "philosophai-papers-complete" More Information needed tabular1K<n<10K0 likes25 downloads1y agoHugging Face12TajikNLPWorld /tajik-news-multilabelgated Dataset Card for Tajik News Multilabel Classification Dataset Details Dataset Description This dataset contains 108,947 Tajik news articles annotated with multiple topic labels. Each document can belong to any subset of 14 predefined tags. The average number of labels per document is 5.27. The labels were generated using a keyword‑based rule system that scans the article text for relevant terms. The dataset is intended for multilabel classification… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-news-multilabel.tabulartext-classification100K<n<1M0 likes25 downloads25d agoHugging Face13halilibr /dilbazlar-anxiety-disorders-recognition-not-augmented-not-anxiety-multilabel-tr-datasettabular10K<n<100K0 likes23 downloads2y agoHugging Face14tshasan /multi-label-web-categorization Multi-Label Web Page Classification Dataset Dataset Description The Multi-Label Web Page Classification Dataset is a curated dataset containingweb page titles and snippets, extracted from the CC-Meta25-1M dataset. Each entry has been automatically categorized into multiple predefined categories using ChatGPT-4o-mini. This dataset is designed for multi-label text classification tasks, making it ideal for training and evaluating machine learning models in web content… See the full description on the dataset page: https://huggingface.co/datasets/tshasan/multi-label-web-categorization.tabulartext-classification10K<n<100K1 likes23 downloads1y agoHugging Face15victoriadreis /TuPY_dataset_multilabel Portuguese Hate Speech Dataset (TuPy) The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML) and natural language processing (NLP) techniques. TuPy is formed by 10000 thousand unpublished annotated tweets collected in 2023. This repository is organized as follows: root. ├── annotations : classification given by annotators ├── raw corpus : dataset before… See the full description on the dataset page: https://huggingface.co/datasets/victoriadreis/TuPY_dataset_multilabel.tabulartext-classification10K<n<100K3 likes21 downloads3y agoHugging Face16BashkirNLPWorld /bashkir-news-multilabelgated Dataset Card for Bashkir News Multilabel Classification Dataset Dataset Details Dataset Description This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.tabulartext-classification10K<n<100K0 likes20 downloads25d agoHugging Face17higopires /RePro-categories-multilabel RePro: A Benchmark Dataset for Opinion Mining in Brazilian Portuguese RePro, which stands for "REview of PROducts," is a benchmark dataset for opinion mining in Brazilian Portuguese. It consists of 10,000 humanly annotated e-commerce product reviews, each labeled with sentiment and topic information. The dataset was created based on data from one of the largest Brazilian e-commerce platforms, which produced the B2W-Reviews01 dataset… See the full description on the dataset page: https://huggingface.co/datasets/higopires/RePro-categories-multilabel.tabulartext-classification10K<n<100K0 likes19 downloads2y agoHugging Face18qquarkq /russian-toxic-multilabel-comments Dataset Card for Toxic Russian Multilabel Comments Описание Этот датасет содержит размеченные комментарии на русском языке по трём независимым бинарным категориям: Profanity (ненормативная лексика) Threat (угрозы) Illegal (нарушение закона) Каждый текст может относиться сразу к нескольким категориям одновременно, либо же ни к одной (нетоксичный текст). Languages Только русский язык (ru). Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/qquarkq/russian-toxic-multilabel-comments.tabulartext-classification10K<n<100K1 likes19 downloads2mo agoHugging Face19Neha2607 /jigsaw-toxic-multilabeltabular100K<n<1M0 likes17 downloads7mo agoHugging Face20nhantran4425 /vnexpress-news-multilabel-2025 VnExpress News Multi-label Dataset 2025 Giới thiệu Bộ dữ liệu ~18,500 bài báo từ VnExpress.net, được gán nhãn đa nhãn với 88 chủ đề. Phù hợp cho bài toán phân loại văn bản tiếng Việt (Vietnamese text classification). Thống kê Train: 14,860 bài Test: 3,715 bài Số nhãn: 88 Ngôn ngữ: Tiếng Việt Tiền xử lý Word segmentation: underthesea Stopwords removal One-hot encoding nhãn Cấu trúc content_final: title×3 + description×2 + content (đã… See the full description on the dataset page: https://huggingface.co/datasets/nhantran4425/vnexpress-news-multilabel-2025.tabulartext-classification10K<n<100K0 likes17 downloads5mo agoHugging Face21sumaiya-afroze /Multi-Label_Bangla_Hate_Speech_Datareadme_text = """ Bangla Hate Speech Extended Dataset 📖 Overview This dataset is an expanded version of the original Bengali Hate Speech Dataset created by Hriteshwar Talukder and Md Saiful Islam. The original dataset provided a strong foundation for hate speech detection in the Bengali language. In this extended version, the dataset has been: Expanded in size with ~5000 additional Bengali social media comments. Reclassified with fine-grained categories… See the full description on the dataset page: https://huggingface.co/datasets/sumaiya-afroze/Multi-Label_Bangla_Hate_Speech_Data.tabulartext-classification10K<n<100K0 likes16 downloads11mo agoHugging Face22hojzas /setfit-proj8-multilabel_2tabularn<1K0 likes15 downloads3y agoHugging Face23imomayiz /multilabel-classification-undersampled-300tabular1K<n<10K0 likes15 downloads1y agoHugging Face24hojzas /setfit-proj8-multilabel_2_validationtabularn<1K0 likes11 downloads3y agoHugging Face25fahrendrakhoirul /ecommerce-reviews-multilabel-dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [Fahrendra K I] Language(s) (NLP): [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/fahrendrakhoirul/ecommerce-reviews-multilabel-dataset.tabulartext-classification1K<n<10K0 likes10 downloads2y agoHugging Face26Sharath45 /mentalhealth_multilabel_classification Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [Sharath Ragav] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Sharath45/mentalhealth_multilabel_classification.tabular10K<n<100K0 likes10 downloads2y agoHugging Face27Alok64 /multilabel_finance_email_inquiriestabular1K<n<10K0 likes9 downloads2y agoHugging Face28JanviC /emotion_multilabel_datasettabular1K<n<10K0 likes8 downloads8mo agoHugging Face29JanviC /stack_multilabel_subset_chattabular1K<n<10K0 likes8 downloads8mo agoHugging Face30hojzas /proj8-multilabeltabularn<1K0 likes7 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.