datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ade_corpus_v2_classification
ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data.
This is a dataset for classification if a sentence is ADE-related (True) or not (False).
Train size: 17,637
Test size: 5,879
Source dataset
Paper
en-document-classification
English Document Classification Dataset
This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora.
Dataset Summary
The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.kinopoisk-sentiment-classificationru-scibench-grnti-classificationmultilingual-safety-classification-dataset
Multilingual Safety Classification Dataset
A multilingual dataset for safety classification across 60 languages, created by Hasan Kurşun through machine translation of English safety prompts using NLLB-200-3.3B.
Dataset Details
Processed by: Hasan KurşunAuthor: Hasan KurşunYear: 2025Source Dataset: mvrcii/safety-moderation-benchmarkTranslation Model: facebook/nllb-200-3.3B
Languages (60)
African Languages (16): Amharic, Hausa, Kinyarwanda, Luganda… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/multilingual-safety-classification-dataset.adaption-defi-wallet-risk-classification
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-defi_wallet_risk_classification
This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks. Each entry provides behavioral features such as transaction counts, action ratios, and concentration metrics within a specific feature window to predict a binary risk label. The completions offer a concise justification for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification.headline-classificationru-reviews-classificationru-scibench-oecd-classificationmulti-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.inappropriateness-classificationLLMTrace_classification
LLMTrace - Classification Dataset
🌐 LLMTrace Website |
📜 LLMTrace Paper on arXiv |
🤗 LLMTrace - Detection Dataset |
🤗 GigaCheck classification model |
This repository contains the Classification portion of the LLMTrace project. This dataset is specifically designed for the binary classification of texts as either human-written or AI-generated.
For full details on the data collection methodology, statistics, and experiments, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/LLMTrace_classification.en-document-format-classification
English Document Format Classification Dataset
English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats.
Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks.
Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage).
Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.en-document-topic-classification
English Document Topic Classification Dataset
English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics.
Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks.
Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation).
Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.slop-classification
Slop classifier dataset
A human-annotated dataset for studying and classifying AI-generated text that people perceive as “AI slop.”
The dataset is built from samples collected from existing public datasets and annotated through the Bench Labs SlopFinder interface.
Slop score
Each sample receives a score based on human votes:
-1 = definitely slop
0 = undecided / neutral
+1 = not slop at all
The score represents human judgment, not an objective measure of quality… See the full description on the dataset page: https://huggingface.co/datasets/bench-labs/slop-classification.wos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification.
The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.astro-classification-redshifts
AstroClassification and Redshifts Datasets
This dataset was used for the AstroClassification and Redshifts introduced in Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations. This is a dataset of simulated astronomical time-series (e.g., supernovae, active galactic nuclei), and the task is to classify the object type (AstroClassification) or predict the object's redshift (Redshifts).
Repository: https://github.com/helenqu/connect-later
Paper: will be… See the full description on the dataset page: https://huggingface.co/datasets/helenqu/astro-classification-redshifts.tlink-classificationastro-classification-redshifts-augmented
AstroClassification and Redshifts Augmented Dataset
This dataset was used in the fine-tuning with targeted augmentations step for the AstroClassification and Redshifts tasks introduced in Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations. This is a dataset of simulated astronomical time-series (e.g., supernovae, active galactic nuclei) augmented with the redshifting targeted augmentation, and the task is to classify the object type… See the full description on the dataset page: https://huggingface.co/datasets/helenqu/astro-classification-redshifts-augmented.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.sensitive-topics-classificationrobotics-arxiv-sustainability-classification
ArXiv Robotics Sustainability Dataset
This dataset compiles sustainability annotations for robotics papers from arXiv metadata files in this repository.
Dataset Structure
Each record is based on an input paper object, with the following transformation rules:
id - url to the paper on arXiv
title - title of the paper
authors - list of author names
published - publication date in ISO format (YYYY-MM-DD)
Added classification fields:
sdg_analysis - original text from… See the full description on the dataset page: https://huggingface.co/datasets/sustainable-robotics/robotics-arxiv-sustainability-classification.heuristic_classification-filtered-pile-50M
Dataset Card for heuristic_classification-filtered-pile-50M
Dataset Summary
This dataset is a subset of The Pile, selected via the heuristic classification data selection method. The target distribution for heuristic classification are the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (51.2M examples) in jsonl format.
Data Instances
{"contents": "Members join for free and… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/heuristic_classification-filtered-pile-50M.Chinese-Legal-Case-Classification-DatasetWikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.shell-safety-v2-classificationConverted from tomngdev/shell-safety-v2 into Text Classification dataset.
Structure is for my own training with static system prompt and changing <SessionContext></SessionContext> block
ctmatch_classificationCTMatch Classification Dataset
This is a combined set of 2 labelled datasets of:
topic (patient descriptions), doc (clinical trials documents - selected fields), and label ({0, 1, 2}) triples, in jsonl format.
(Somewhat of a duplication of some of the ir_dataset also available on HF.)
These have been processed using ctproc, and in this state can be used by various tokenizers for fine-tuning (see ctmatch for examples).
These 2 datasets contain no patient identifying information are openly… See the full description on the dataset page: https://huggingface.co/datasets/semaj83/ctmatch_classification.multiclass-email-classificationThis dataset comprises of more than 2000 emails across multiple categories, which can he helpful for tasks like LLM training and fine-tuning. The dataset is also provided with a python script that would generate emails automatically
The dataset contains email across 10 different categories namely, "Business", "Personal", "Promotions", "Customer Support", "Job Application", "Finance & Bills", "Events & Invitations", "Travel & Bookings", "Reminders", "Newsletters"
Total emails: 2105
Label… See the full description on the dataset page: https://huggingface.co/datasets/imnim/multiclass-email-classification.cosmopedia-classificationcedr-classification
