datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.short-text-multi-labeled-emotion-classificationmultilabel-tagalog-hate-speechcourse-review-multilabel-sentiment-analysisTuPY_dataset_multilabel
Portuguese Hate Speech Dataset (TuPy)
The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML) and natural language processing (NLP) techniques. TuPy is formed by 10000 thousand unpublished annotated tweets collected in 2023.
This repository is organized as follows:
root.
├── annotations : classification given by annotators
├── raw corpus : dataset before… See the full description on the dataset page: https://huggingface.co/datasets/victoriadreis/TuPY_dataset_multilabel.vnexpress-news-multilabel-2025
VnExpress News Multi-label Dataset 2025
Giới thiệu
Bộ dữ liệu ~18,500 bài báo từ VnExpress.net, được gán nhãn đa nhãn với 88 chủ đề.
Phù hợp cho bài toán phân loại văn bản tiếng Việt (Vietnamese text classification).
Thống kê
Train: 14,860 bài
Test: 3,715 bài
Số nhãn: 88
Ngôn ngữ: Tiếng Việt
Tiền xử lý
Word segmentation: underthesea
Stopwords removal
One-hot encoding nhãn
Cấu trúc
content_final: title×3 + description×2 + content (đã… See the full description on the dataset page: https://huggingface.co/datasets/nhantran4425/vnexpress-news-multilabel-2025.Multi-Label_Bangla_Hate_Speech_Datareadme_text = """
Bangla Hate Speech Extended Dataset
📖 Overview
This dataset is an expanded version of the original Bengali Hate Speech Dataset created by Hriteshwar Talukder and Md Saiful Islam.
The original dataset provided a strong foundation for hate speech detection in the Bengali language. In this extended version, the dataset has been:
Expanded in size with ~5000 additional Bengali social media comments.
Reclassified with fine-grained categories… See the full description on the dataset page: https://huggingface.co/datasets/sumaiya-afroze/Multi-Label_Bangla_Hate_Speech_Data.setfit-proj8-multilabel_2setfit-proj8-multilabel_2_validationmentalhealth_multilabel_classification
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Sharath Ragav]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Sharath45/mentalhealth_multilabel_classification.eyeR-classification-multi-label-category2multilabel_finance_email_inquiriesproj8-multilabelproj8-multilabel-validationmultilabel-classificationeyeR-classification-multi-label-category1Multilabel_Emotionmultilabel-sociotechnical_imaginaries-2025_05_06multi-label-review-1000stack_multilabel_subsetrg-7wildlife-multilabel-v1
