datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification.
The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.tram-attack-multilabel-clean
TRAM ATT&CK Multi-Label (cleaned, with leak-free splits)
Sentence-level multi-label mapping of cyber threat intelligence prose to MITRE
ATT&CK technique IDs. Derived from MITRE CTID's
TRAM corpus,
deduplicated and republished with two split schemes so that leakage can be
measured rather than assumed.
Everything here is regenerated by python scripts/01_build_dataset.py. No row
was edited by hand.
Why this exists
The upstream corpus contains 19,178 sentences drawn… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/tram-attack-multilabel-clean.MADE-Multilabel-Benchmark
MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events (ACL 2026)
Authors: Raunak Agarwal, Markus Wenzel, Simon Baur, Jonas Zimmer, George Harvey, Jackie Ma
Blog; Project Page; Github; Arxiv
Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification (UQ) to support human oversight.
Multi-label text… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/MADE-Multilabel-Benchmark.medical-emails-multilabel-dataset
Medical Emails Multi-Label Classification Dataset
This dataset contains 1000 synthetic medical emails for multi-label classification across 5 combined categories.
Categories (200 emails each)
Category
Count
Description
Adverse Event + Medical Information
200
Emails reporting adverse symptoms while also requesting clinical/scientific information
Medical Information + Product Complaint
200
Emails requesting clinical guidance while also reporting product quality… See the full description on the dataset page: https://huggingface.co/datasets/Ramesh10/medical-emails-multilabel-dataset.CEN_multilabelmedical-emails-multi-label-dataset
Ramesh10/medical-emails-multi-label-dataset
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('Ramesh10/medical-emails-multi-label-dataset')
ecommerce-reviews-multilabel-dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Fahrendra K I]
Language(s) (NLP): [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/fahrendrakhoirul/ecommerce-reviews-multilabel-dataset.multi_label_datasethierarchical_multi_label_dataset
