datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
research_titles_multi-labelPubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.arxiv-abstract-multilabelwos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification.
The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.gklmip-news-khm-multilabelclassification
GKLMIPNews_khm_MultiLabelClassification
Deduplicated copy of kornwtp/gklmip-news-khm-multilabelclassification.
Splits
split
rows
test
1,398
train
3,959
validation
1,376
vlsp2018sa-hotel-vie-multilabelclassification
VLSP2018SAHotel_vie_MultiLabelClassification
Deduplicated copy of kornwtp/vlsp2018sa-hotel-vie-multilabelclassification.
Splits
split
rows
test
599
train
2,947
validation
1,985
hatespeech-ind-multilabelclassification
HateSpeech_ind_MultiLabelClassification
Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification.
Splits
split
rows
train
13,014
netifier-ind-multilabelclassification
Netifier_ind_MultiLabelClassification
Deduplicated copy of kornwtp/netifier-ind-multilabelclassification.
Splits
split
rows
test
750
train
6,361
russian-toxic-comments-multilabel
Russian Toxic Comments Multi-label Dataset
Dataset Description
Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности.
Цель
Обучение модели для автоматического обнаружения трех типов токсичного контента:
Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань
Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.prachathai67k-tha-multilabelclassification
Prachathai67k_tha_MultiLabelClassification
Deduplicated copy of kornwtp/prachathai67k-tha-multilabelclassification.
Splits
split
rows
train
67,488
wds_voc2007_multilabelcasa-ind-multilabelclassification
CASA_ind_MultiLabelClassification
Deduplicated copy of kornwtp/casa-ind-multilabelclassification.
Splits
split
rows
test
180
train
809
validation
90
dengue-fil-multilabelclassification
Dengue_fil_MultiLabelClassification
Deduplicated copy of kornwtp/dengue-fil-multilabelclassification.
Splits
split
rows
test
494
train
3,924
validation
498
vlsp2018sa-restaurant-vie-multilabelclassification
VLSP2018SARestaurant_vie_MultiLabelClassification
Deduplicated copy of kornwtp/vlsp2018sa-restaurant-vie-multilabelclassification.
Splits
split
rows
test
499
train
2,958
validation
1,289
prachathai67k-tha-multilabelclassificationref: https://github.com/PyThaiNLP/prachathai-67k
icd10cm-multilabel-promptawesome-japanese-nlp-multilabel-dataset
Dataset overview
This is a dataset for Japanese natural language processing with multi-label annotations of research field labels for GitHub repositories in the NLP domain.
Please refer to this paper for the specific method of constructing the dataset. It is written in Japanese.
Input and Output
Input: Information from GitHub repositories (description, README text, PDF text, screenshot images)
Output: Multi-label classification of NLP research fields
Problem Setting of the… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/awesome-japanese-nlp-multilabel-dataset.hoasa-ind-multilabelclassification
HoASA_ind_MultiLabelClassification
Deduplicated copy of kornwtp/hoasa-ind-multilabelclassification.
Splits
split
rows
test
286
train
2,267
validation
285
scikit-learn-issues-multilabel
🧩 Scikit-learn GitHub Issues – Multilabel Dataset
This dataset contains GitHub issues from the scikit-learn repository, prepared for multilabel NLP tasks such as issue tagging, automated triage, and semantic search.
Each row corresponds to one issue-comment context, making the dataset suitable for real-world developer tooling.
📌 Motivation
GitHub issues are a critical signal in open-source projects:
Bug tracking
Feature requests
Documentation improvements… See the full description on the dataset page: https://huggingface.co/datasets/Talip7/scikit-learn-issues-multilabel.multi-label-food-recognition
Multi-Label Food Recognition Dataset
This is a multi-label food recognition dataset generated from single-class food images.
Each image contains 2-5 different food items composited together using natural composition methods.
Dataset Details
Total Images: 13,000
Training Images: 10,400 (80%)
Validation Images: 2,600 (20%)
Number of Classes: 90
Labels per Image: 2-5 labels
Image Format: RGB, 512x512 pixels
File Format: Parquet
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.prachathai67k-mya-multilabelclassification
Prachathai67k_mya_MultiLabelClassification
Deduplicated copy of kornwtp/prachathai67k-mya-multilabelclassification.
Splits
split
rows
train
2,153
truevoice-intent-tha-multilabelclassification
TrueVoiceIntent_tha_MultiLabelClassification
Deduplicated copy of kornwtp/truevoice-intent-tha-multilabelclassification.
Splits
split
rows
train
13,355
toxicity-multi-label-classifier
Part of a course titled "Generative AI application design & development"
https://genai.acloudfan.com/
Created from a dataset available on Kaggle.
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data
synthetic-text-classification-news-multi-label
Dataset Card for synthetic-text-classification-news-multi-label
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/synthetic-text-classification-news-multi-label/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news-multi-label.environment-multi-labels-even
Dataset Card for "environment-multi-labels-even"
More Information needed
multilabel-imagenet-1k
MultiLabel ImageNet-1K Train Annotations with Selected Masks
This dataset contains automated multi-label annotations for the ImageNet-1K training split, together with spatial masks for the selected object-level labels.
The release is designed to make the annotations easy to inspect and reuse. It does not include the original ImageNet images. Users need access to the ImageNet-1K training images separately; image paths are stored relative to the ImageNet train root, for example:… See the full description on the dataset page: https://huggingface.co/datasets/k3999/multilabel-imagenet-1k.toxic-russian-comments-multilabel
Russian Toxic Comments Multi-Label Balanced Dataset
Описание
Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками:
profanity: наличие нецензурной лексики (мат)
threat: наличие угроз
illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT)
Структура данных
Датасет содержит следующие поля:
Поле
Тип
Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.short-text-multi-labeled-emotion-classificationvlsp2018sa-restaurant-vie-multilabelclassificationref: https://github.com/vndee/awsome-vietnamese-nlp
