datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.arxiv-abstract-multilabelrussian-toxic-comments-multilabel
Russian Toxic Comments Multi-label Dataset
Dataset Description
Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности.
Цель
Обучение модели для автоматического обнаружения трех типов токсичного контента:
Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань
Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.toxic-russian-comments-multilabel
Russian Toxic Comments Multi-Label Balanced Dataset
Описание
Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками:
profanity: наличие нецензурной лексики (мат)
threat: наличие угроз
illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT)
Структура данных
Датасет содержит следующие поля:
Поле
Тип
Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.hoasa-ind-multilabelclassificationdengue-fil-multilabelclassificationmultilabel-tagalog-hate-speechtatar-news-analysis-multilabel
Dataset Card for Tatar News Multilabel Classification
Dataset Details
Dataset Description
The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.course-review-multilabel-sentiment-analysiscasa-ind-multilabelclassificationphilosophy-schools-multilabel
Dataset Card for "philosophai-papers-complete"
More Information needed
tajik-news-multilabel
Dataset Card for Tajik News Multilabel Classification
Dataset Details
Dataset Description
This dataset contains 108,947 Tajik news articles annotated with multiple topic labels. Each document can belong to any subset of 14 predefined tags. The average number of labels per document is 5.27. The labels were generated using a keyword‑based rule system that scans the article text for relevant terms. The dataset is intended for multilabel classification… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-news-multilabel.dilbazlar-anxiety-disorders-recognition-not-augmented-not-anxiety-multilabel-tr-datasetmulti-label-web-categorization
Multi-Label Web Page Classification Dataset
Dataset Description
The Multi-Label Web Page Classification Dataset is a curated dataset containingweb page titles and snippets, extracted from the CC-Meta25-1M dataset. Each entry has been automatically categorized into multiple predefined categories using ChatGPT-4o-mini.
This dataset is designed for multi-label text classification tasks, making it ideal for training and evaluating machine learning models in web content… See the full description on the dataset page: https://huggingface.co/datasets/tshasan/multi-label-web-categorization.TuPY_dataset_multilabel
Portuguese Hate Speech Dataset (TuPy)
The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML) and natural language processing (NLP) techniques. TuPy is formed by 10000 thousand unpublished annotated tweets collected in 2023.
This repository is organized as follows:
root.
├── annotations : classification given by annotators
├── raw corpus : dataset before… See the full description on the dataset page: https://huggingface.co/datasets/victoriadreis/TuPY_dataset_multilabel.bashkir-news-multilabel
Dataset Card for Bashkir News Multilabel Classification Dataset
Dataset Details
Dataset Description
This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.RePro-categories-multilabel
RePro: A Benchmark Dataset for Opinion Mining in Brazilian Portuguese
RePro, which stands for "REview of PROducts," is a benchmark dataset for opinion mining in Brazilian Portuguese. It consists of 10,000 humanly annotated e-commerce product reviews, each labeled with sentiment and topic information. The dataset was created based on data from one of the largest Brazilian e-commerce platforms, which produced the B2W-Reviews01 dataset… See the full description on the dataset page: https://huggingface.co/datasets/higopires/RePro-categories-multilabel.russian-toxic-multilabel-comments
Dataset Card for Toxic Russian Multilabel Comments
Описание
Этот датасет содержит размеченные комментарии на русском языке по трём независимым бинарным категориям:
Profanity (ненормативная лексика)
Threat (угрозы)
Illegal (нарушение закона)
Каждый текст может относиться сразу к нескольким категориям одновременно, либо же ни к одной (нетоксичный текст).
Languages
Только русский язык (ru).
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/qquarkq/russian-toxic-multilabel-comments.jigsaw-toxic-multilabelvnexpress-news-multilabel-2025
VnExpress News Multi-label Dataset 2025
Giới thiệu
Bộ dữ liệu ~18,500 bài báo từ VnExpress.net, được gán nhãn đa nhãn với 88 chủ đề.
Phù hợp cho bài toán phân loại văn bản tiếng Việt (Vietnamese text classification).
Thống kê
Train: 14,860 bài
Test: 3,715 bài
Số nhãn: 88
Ngôn ngữ: Tiếng Việt
Tiền xử lý
Word segmentation: underthesea
Stopwords removal
One-hot encoding nhãn
Cấu trúc
content_final: title×3 + description×2 + content (đã… See the full description on the dataset page: https://huggingface.co/datasets/nhantran4425/vnexpress-news-multilabel-2025.Multi-Label_Bangla_Hate_Speech_Datareadme_text = """
Bangla Hate Speech Extended Dataset
📖 Overview
This dataset is an expanded version of the original Bengali Hate Speech Dataset created by Hriteshwar Talukder and Md Saiful Islam.
The original dataset provided a strong foundation for hate speech detection in the Bengali language. In this extended version, the dataset has been:
Expanded in size with ~5000 additional Bengali social media comments.
Reclassified with fine-grained categories… See the full description on the dataset page: https://huggingface.co/datasets/sumaiya-afroze/Multi-Label_Bangla_Hate_Speech_Data.setfit-proj8-multilabel_2multilabel-classification-undersampled-300setfit-proj8-multilabel_2_validationecommerce-reviews-multilabel-dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Fahrendra K I]
Language(s) (NLP): [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/fahrendrakhoirul/ecommerce-reviews-multilabel-dataset.mentalhealth_multilabel_classification
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Sharath Ragav]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Sharath45/mentalhealth_multilabel_classification.multilabel_finance_email_inquiriesemotion_multilabel_datasetstack_multilabel_subset_chatproj8-multilabel
