datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.short-text-multi-labeled-emotion-classificationmultilabel-tagalog-hate-speechcourse-review-multilabel-sentiment-analysisMulti-label-Prompt-Dataset
Multi-label Prompt Dataset
A multi-label prompt classification corpus designed for training lightweight, CPU-efficient machine learning models (such as CatBoost, LightGBM, and FastText) for prompt complexity estimation, task intent classification, output token length forecasting, and dynamic LLM routing.
Dataset Summary
Total Unique Samples: 1,859 deduplicated prompts
Number of Classes: 23 multi-label tags across 4 semantic dimensions
Language: English (en)… See the full description on the dataset page: https://huggingface.co/datasets/Nasim435/Multi-label-Prompt-Dataset.TuPY_dataset_multilabel
Portuguese Hate Speech Dataset (TuPy)
The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML) and natural language processing (NLP) techniques. TuPy is formed by 10000 thousand unpublished annotated tweets collected in 2023.
This repository is organized as follows:
root.
├── annotations : classification given by annotators
├── raw corpus : dataset before… See the full description on the dataset page: https://huggingface.co/datasets/victoriadreis/TuPY_dataset_multilabel.multi-label-text-classificationsetfit-proj8-multilabel_2vnexpress-news-multilabel-2025
VnExpress News Multi-label Dataset 2025
Giới thiệu
Bộ dữ liệu ~18,500 bài báo từ VnExpress.net, được gán nhãn đa nhãn với 88 chủ đề.
Phù hợp cho bài toán phân loại văn bản tiếng Việt (Vietnamese text classification).
Thống kê
Train: 14,860 bài
Test: 3,715 bài
Số nhãn: 88
Ngôn ngữ: Tiếng Việt
Tiền xử lý
Word segmentation: underthesea
Stopwords removal
One-hot encoding nhãn
Cấu trúc
content_final: title×3 + description×2 + content (đã… See the full description on the dataset page: https://huggingface.co/datasets/nhantran4425/vnexpress-news-multilabel-2025.risk_sig_train_multilabel_OPRMulti-Label_Bangla_Hate_Speech_Datareadme_text = """
Bangla Hate Speech Extended Dataset
📖 Overview
This dataset is an expanded version of the original Bengali Hate Speech Dataset created by Hriteshwar Talukder and Md Saiful Islam.
The original dataset provided a strong foundation for hate speech detection in the Bengali language. In this extended version, the dataset has been:
Expanded in size with ~5000 additional Bengali social media comments.
Reclassified with fine-grained categories… See the full description on the dataset page: https://huggingface.co/datasets/sumaiya-afroze/Multi-Label_Bangla_Hate_Speech_Data.Bangla-Multi-Label-Text_AnalysisThis dataset is a consolidated collection of four public Bengali text datasets curated for sentiment analysis, toxic comment classification and bengali news classification. It consists of Bengali text comments annotated with multiple categories, covering a wide range of sentiment and content-based labels. Aimed at advancing research in Bengali language processing, this dataset is particularly suited for tasks like sentiment analysis, hate speech detection, and contextual comment… See the full description on the dataset page: https://huggingface.co/datasets/raselmeya2194/Bangla-Multi-Label-Text_Analysis.risk_sig_train_multilabel_FIN_25ksetfit-proj8-multilabel_2_validationmentalhealth_multilabel_classification
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Sharath Ragav]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Sharath45/mentalhealth_multilabel_classification.AURA-Multi_Label_Classification
AURA-Classification (Multi-Label Version)
Dataset Description
The AURA (App User Review in Arabic) Classification dataset is a collection of 2,900 Arabic-language app reviews collected from various mobile applications. This dataset is designed for multi-label text classification, where each review can belong to multiple classes simultaneously.
Each review in the dataset was independently annotated by five different annotators. To construct the multi-label version of the… See the full description on the dataset page: https://huggingface.co/datasets/irfan-ahmad/AURA-Multi_Label_Classification.multilabel_finance_email_inquiriesproj8-multilabeleyeR-classification-multi-label-category2Multilabel_Emotionproj8-multilabel-validationmultilabel-classificationeyeR-classification-multi-label-category1multilabel-sociotechnical_imaginaries-2025_05_06multi-label-review-1000stack_multilabel_subsetrg-7wildlife-multilabel-v1
