datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.arxiv-abstract-multilabelAID_MultiLabel
Dataset Card for "AID_MultiLabel"
Licensing Information
CC0: Public Domain
Citation Information
Imagery:
AID: A benchmark data set for performance evaluation of aerial scene classification
Multilabels:
Relation Network for Multi-label Aerial Image Classification
@article{xia2017aid,
title = {AID: A benchmark data set for performance evaluation of aerial scene classification},
author = {Xia, Gui-Song and Hu, Jingwen and Hu, Fan and Shi, Baoguang… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/AID_MultiLabel.id_multilabel_hsThe ID_MULTILABEL_HS dataset is collection of 13,169 tweets in Indonesian language,
designed for hate speech detection NLP task. This dataset is combination from previous research and newly crawled data from Twitter.
This is a multilabel dataset with label details as follows:
-HS : hate speech label;
-Abusive : abusive language label;
-HS_Individual : hate speech targeted to an individual;
-HS_Group : hate speech targeted to a group;
-HS_Religion : hate speech related to religion/creed;
-HS_Race : hate speech related to race/ethnicity;
-HS_Physical : hate speech related to physical/disability;
-HS_Gender : hate speech related to gender/sexual orientation;
-HS_Gender : hate related to other invective/slander;
-HS_Weak : weak hate speech;
-HS_Moderate : moderate hate speech;
-HS_Strong : strong hate speech.gklmip-news-khm-multilabelclassification
GKLMIPNews_khm_MultiLabelClassification
Deduplicated copy of kornwtp/gklmip-news-khm-multilabelclassification.
Splits
split
rows
test
1,398
train
3,959
validation
1,376
vlsp2018sa-hotel-vie-multilabelclassification
VLSP2018SAHotel_vie_MultiLabelClassification
Deduplicated copy of kornwtp/vlsp2018sa-hotel-vie-multilabelclassification.
Splits
split
rows
test
599
train
2,947
validation
1,985
hatespeech-ind-multilabelclassification
HateSpeech_ind_MultiLabelClassification
Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification.
Splits
split
rows
train
13,014
netifier-ind-multilabelclassification
Netifier_ind_MultiLabelClassification
Deduplicated copy of kornwtp/netifier-ind-multilabelclassification.
Splits
split
rows
test
750
train
6,361
Multilabel-Portrait-18K
Multilabel-Portrait-18K
Multilabel-Portrait-18K is a multi-label portrait classification dataset designed to analyze and categorize different styles of portrait images. It supports classification into the following four portrait types:
0 — Anime Portrait
1 — Cartoon Portrait
2 — Real Portrait
3 — Sketch Portrait
This dataset is ideal for training and evaluating machine learning models in the domain of portrait-style classification. The goal is to enable accurate recognition… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Multilabel-Portrait-18K.russian-toxic-comments-multilabel
Russian Toxic Comments Multi-label Dataset
Dataset Description
Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности.
Цель
Обучение модели для автоматического обнаружения трех типов токсичного контента:
Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань
Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.prachathai67k-tha-multilabelclassification
Prachathai67k_tha_MultiLabelClassification
Deduplicated copy of kornwtp/prachathai67k-tha-multilabelclassification.
Splits
split
rows
train
67,488
UC_Merced_LandUse_MultiLabel
Dataset Card for "UC_Merced_LandUse_MultiLabel"
Licensing Information
Public Domain; “Map services and data available from U.S. Geological Survey, National Geospatial Program.”
Citation Information
Imagery:
Bag-of-visual-words and spatial extensions for land-use classification
Multilabels:
Multilabel Remote Sensing Image Retrieval Using a Semisupervised Graph-Theoretic Method
@inproceedings{yang2010bag,
title = {Bag-of-visual-words and spatial… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/UC_Merced_LandUse_MultiLabel.wds_voc2007_multilabelcasa-ind-multilabelclassification
CASA_ind_MultiLabelClassification
Deduplicated copy of kornwtp/casa-ind-multilabelclassification.
Splits
split
rows
test
180
train
809
validation
90
dengue-fil-multilabelclassification
Dengue_fil_MultiLabelClassification
Deduplicated copy of kornwtp/dengue-fil-multilabelclassification.
Splits
split
rows
test
494
train
3,924
validation
498
vlsp2018sa-restaurant-vie-multilabelclassification
VLSP2018SARestaurant_vie_MultiLabelClassification
Deduplicated copy of kornwtp/vlsp2018sa-restaurant-vie-multilabelclassification.
Splits
split
rows
test
499
train
2,958
validation
1,289
prachathai67k-tha-multilabelclassificationref: https://github.com/PyThaiNLP/prachathai-67k
icd10cm-multilabel-promptawesome-japanese-nlp-multilabel-dataset
Dataset overview
This is a dataset for Japanese natural language processing with multi-label annotations of research field labels for GitHub repositories in the NLP domain.
Please refer to this paper for the specific method of constructing the dataset. It is written in Japanese.
Input and Output
Input: Information from GitHub repositories (description, README text, PDF text, screenshot images)
Output: Multi-label classification of NLP research fields
Problem Setting of the… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/awesome-japanese-nlp-multilabel-dataset.hoasa-ind-multilabelclassification
HoASA_ind_MultiLabelClassification
Deduplicated copy of kornwtp/hoasa-ind-multilabelclassification.
Splits
split
rows
test
286
train
2,267
validation
285
scikit-learn-issues-multilabel
🧩 Scikit-learn GitHub Issues – Multilabel Dataset
This dataset contains GitHub issues from the scikit-learn repository, prepared for multilabel NLP tasks such as issue tagging, automated triage, and semantic search.
Each row corresponds to one issue-comment context, making the dataset suitable for real-world developer tooling.
📌 Motivation
GitHub issues are a critical signal in open-source projects:
Bug tracking
Feature requests
Documentation improvements… See the full description on the dataset page: https://huggingface.co/datasets/Talip7/scikit-learn-issues-multilabel.multi-label-food-recognition
Multi-Label Food Recognition Dataset
This is a multi-label food recognition dataset generated from single-class food images.
Each image contains 2-5 different food items composited together using natural composition methods.
Dataset Details
Total Images: 13,000
Training Images: 10,400 (80%)
Validation Images: 2,600 (20%)
Number of Classes: 90
Labels per Image: 2-5 labels
Image Format: RGB, 512x512 pixels
File Format: Parquet
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.prachathai67k-mya-multilabelclassification
Prachathai67k_mya_MultiLabelClassification
Deduplicated copy of kornwtp/prachathai67k-mya-multilabelclassification.
Splits
split
rows
train
2,153
truevoice-intent-tha-multilabelclassification
TrueVoiceIntent_tha_MultiLabelClassification
Deduplicated copy of kornwtp/truevoice-intent-tha-multilabelclassification.
Splits
split
rows
train
13,355
multilabel-imagenet-1k
MultiLabel ImageNet-1K Train Annotations with Selected Masks
This dataset contains automated multi-label annotations for the ImageNet-1K training split, together with spatial masks for the selected object-level labels.
The release is designed to make the annotations easy to inspect and reuse. It does not include the original ImageNet images. Users need access to the ImageNet-1K training images separately; image paths are stored relative to the ImageNet train root, for example:… See the full description on the dataset page: https://huggingface.co/datasets/k3999/multilabel-imagenet-1k.toxic-russian-comments-multilabel
Russian Toxic Comments Multi-Label Balanced Dataset
Описание
Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками:
profanity: наличие нецензурной лексики (мат)
threat: наличие угроз
illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT)
Структура данных
Датасет содержит следующие поля:
Поле
Тип
Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.vlsp2018sa-restaurant-vie-multilabelclassificationref: https://github.com/vndee/awsome-vietnamese-nlp
PubMed-MultiLabel-MeSH
PubMed MultiLabel Text Classification (MeSH)
A dataset of 50,000 PubMed biomedical articles, each manually annotated
by domain experts with MeSH (Medical Subject Headings) labels. With
21,918 unique labels and a mean of ~12.7 labels per document, this is a
densely-labeled extreme multi-label classification benchmark.
Dataset Description
Property
Value
Train examples
40,000
Test examples
10,000
Total unique MeSH labels
21,918
Mean labels per document
~12.7… See the full description on the dataset page: https://huggingface.co/datasets/Tellurio/PubMed-MultiLabel-MeSH.dengue-fil-multilabelclassification
